AnticiPro scored against NASA's own safety analysts
We gave our agentic retrieval platform 11,557 raw aviation incident narratives from NASA's Aviation Safety Reporting System and graded its structured analysis and cited answers against the codes NASA's analysts had already assigned. Every number below is from a frozen protocol on public data that anyone can download and rerun.
The problem we tested
Operational organizations collect free-text reports at volume: after-action write-ups, incident narratives, inspection notes, help-desk tickets. The people who need answers from that material are usually not data scientists. They want to ask a plain question, get an answer, and see the evidence behind it.
Testing a platform on that task is hard because most real corpora have no answer key. ASRS does. Pilots, controllers, mechanics and dispatchers submit narratives, and NASA analysts code each report against controlled vocabularies: the anomaly, the primary problem, contributing factors, human factors and phase of flight. That coding is expert ground truth. We used it as the grader.
What AnticiPro does
AnticiPro is IPNS's agentic retrieval-augmented generation platform for enterprise knowledge, developed at private expense since October 2025. Four stages of its analytic pipeline were under test.
Ingestion-time tagging
Each document is read on arrival and tagged against the controlled vocabularies. The tags become queryable metadata, the machine equivalent of the analyst coding pass.
Query understanding
A plain question such as "which reports involved fatigue while parked?" becomes a typed query with structured constraints, not a keyword search.
Hybrid retrieval
Candidates come from the tag index (exact match on coded facets and dates) and from dense semantic retrieval over the narratives.
Grounded citation
Each candidate is verified against the question using the corpus's own coding conventions, and every answer returns specific reports with a verbatim quotation from each.
Task A · Structured analysis of 1,443 unseen reports
Zero-shot, it matched a classifier trained on 8,000 labelled reports
The comparison that matters is the second number in each tile. A TF-IDF classifier trained on about 8,000 of NASA's labelled reports is a strong supervised baseline for this task. AnticiPro saw no labels at all. It read each narrative with the vocabulary definitions and produced the code, the supporting quotation and a confidence. On the primary problem it tied the supervised model within statistical noise, and on contributing factors it came within two points. The 97% verbatim-span figure is the explainability requirement expressed as a measurement: in 1,401 of 1,443 reports, the quotation the platform cited exists word for word in the narrative.
Task B: answering analytical questions with cited report sets
The second task was harder and we report it as measured. We asked 160 held-out questions of the kind an analyst asks, such as which reports in a given month involved a particular human factor during a particular phase of flight, and scored the platform's cited answer set against the set NASA's coding implies.
| Question family | Answer-set F1 |
|---|---|
| Anomaly type in a given month | 0.308 |
| Human factor by phase of flight | 0.154 |
| Contributing factor by phase of flight | 0.143 |
| Primary problem with a human factor | 0.133 |
| All 160 questions | 0.185 (recall 37.6%, precision 14.6%) |
A keyword baseline (TF-IDF retrieval) scored lower on the same questions. The controlled comparison showed where the ceiling is: when the ingestion tags were of the higher quality the Task A prompt produces, answer-set quality rose on every question family. Tag quality, not retrieval, is the lever, and it is the same lever that drives Task A. We state this because it is what the evaluation found, and because a buyer should know which part of the system improves which number.
Speed and cost
The frozen pipelines were rerun unchanged for timing, from a workstation over the commercial internet, so network time is inside every figure.
| Workload | Users | Median | 90th pct | Throughput |
|---|---|---|---|---|
| Classify one report on ingestion | 1 | 0.93 s | 1.34 s | 61 reports/min |
| Classify, eight in parallel | 8 | 0.93 s | 1.26 s | 472 reports/min |
| Conversational query to a cited answer, served configuration | 1 | 2.0 s | 3.1 s | 16 queries/min |
| Conversational query, sequential configuration | 8 | 8.1 s | 9.1 s | 50 queries/min |
At eight-wide, the full 11,557-report corpus classifies in about 25 minutes. The complete Task A development pass over 10,114 reports cost about $4.20 in model inference; the whole evidence phase ran for under $30. Eight simultaneous users saw no latency degradation in the sequential configuration.
What this means for a buyer
No training data required to start. Zero-shot analysis matched a supervised model on this corpus. A new organization's reports can be coded on day one against its own vocabulary.
Every answer carries its evidence. The verbatim-span rate is measured, not asserted. A reviewer verifies against a quotation rather than approving an opaque output.
Reproducible. The corpus is public, the protocol is frozen, and the baselines are stated. An evaluator can rerun it.
Deployable where the data lives. The model interface is interchangeable, so the same pipeline runs against the model service an authorization boundary allows, including Government-authorized cloud regions, without a redesign.
Limits we state
ASRS is civil aviation safety reporting with NASA's taxonomy and sanitized single-narrative records. What transfers to other domains is the task shape: free-text operator reporting, expert-coded ground truth, and non-specialist users who need answers they can check. Text ingestion was tested; audio and graphics extraction were not. Held-out Task B used a smaller corpus than development, which makes set retrieval easier, and we report both sizes for that reason.
IP Network Solutions, Inc. · Leesburg, Virginia · Federal IT services and AI software since 2003.
Evaluation designed and performed by Michael Young, Program Director, AI Portfolio. Questions and demo requests: MYoung@ipnsinc.com

