IPNS Logo
Case study · IPNS AI Software

AnticiPro scored against NASA's own safety analysts

We gave our agentic retrieval platform 11,557 raw aviation incident narratives from NASA's Aviation Safety Reporting System and graded its structured analysis and cited answers against the codes NASA's analysts had already assigned. Every number below is from a frozen protocol on public data that anyone can download and rerun.

EvaluationAugust 2026CorpusNASA ASRS, publicReports10,114 development · 1,443 held outFundingIPNS internal R&D

The problem we tested

Operational organizations collect free-text reports at volume: after-action write-ups, incident narratives, inspection notes, help-desk tickets. The people who need answers from that material are usually not data scientists. They want to ask a plain question, get an answer, and see the evidence behind it.

Testing a platform on that task is hard because most real corpora have no answer key. ASRS does. Pilots, controllers, mechanics and dispatchers submit narratives, and NASA analysts code each report against controlled vocabularies: the anomaly, the primary problem, contributing factors, human factors and phase of flight. That coding is expert ground truth. We used it as the grader.

What AnticiPro does

AnticiPro is IPNS's agentic retrieval-augmented generation platform for enterprise knowledge, developed at private expense since October 2025. Four stages of its analytic pipeline were under test.

Ingestion-time tagging

Each document is read on arrival and tagged against the controlled vocabularies. The tags become queryable metadata, the machine equivalent of the analyst coding pass.

Query understanding

A plain question such as "which reports involved fatigue while parked?" becomes a typed query with structured constraints, not a keyword search.

Hybrid retrieval

Candidates come from the tag index (exact match on coded facets and dates) and from dense semantic retrieval over the narratives.

Grounded citation

Each candidate is verified against the question using the corpus's own coding conventions, and every answer returns specific reports with a verbatim quotation from each.

System under test. The pipeline is model-agnostic: generation, judgment and embedding go through an interchangeable model interface. These results were measured on one frozen configuration, Vertex AI Gemini 2.5 Flash for generation and judgment and Vertex AI text-embedding-005 for retrieval, so that they are reproducible. The user interface and data connectors were not part of the scored run and no claim here rests on them.

Task A · Structured analysis of 1,443 unseen reports

Zero-shot, it matched a classifier trained on 8,000 labelled reports

68.7%
Primary problem, exact match to NASA's code
majority guess 35.9% · trained TF-IDF 69.5%
0.661
Contributing factors, micro-F1
trained TF-IDF 0.680
0.490
Anomaly type, micro-F1 across 66 labels
majority guess 0.247
97%
Cited evidence spans found verbatim in the source narrative
1,401 of 1,443 reports

The comparison that matters is the second number in each tile. A TF-IDF classifier trained on about 8,000 of NASA's labelled reports is a strong supervised baseline for this task. AnticiPro saw no labels at all. It read each narrative with the vocabulary definitions and produced the code, the supporting quotation and a confidence. On the primary problem it tied the supervised model within statistical noise, and on contributing factors it came within two points. The 97% verbatim-span figure is the explainability requirement expressed as a measurement: in 1,401 of 1,443 reports, the quotation the platform cited exists word for word in the narrative.

Task B: answering analytical questions with cited report sets

The second task was harder and we report it as measured. We asked 160 held-out questions of the kind an analyst asks, such as which reports in a given month involved a particular human factor during a particular phase of flight, and scored the platform's cited answer set against the set NASA's coding implies.

Question familyAnswer-set F1
Anomaly type in a given month0.308
Human factor by phase of flight0.154
Contributing factor by phase of flight0.143
Primary problem with a human factor0.133
All 160 questions0.185 (recall 37.6%, precision 14.6%)

A keyword baseline (TF-IDF retrieval) scored lower on the same questions. The controlled comparison showed where the ceiling is: when the ingestion tags were of the higher quality the Task A prompt produces, answer-set quality rose on every question family. Tag quality, not retrieval, is the lever, and it is the same lever that drives Task A. We state this because it is what the evaluation found, and because a buyer should know which part of the system improves which number.

Speed and cost

The frozen pipelines were rerun unchanged for timing, from a workstation over the commercial internet, so network time is inside every figure.

WorkloadUsersMedian90th pctThroughput
Classify one report on ingestion10.93 s1.34 s61 reports/min
Classify, eight in parallel80.93 s1.26 s472 reports/min
Conversational query to a cited answer, served configuration12.0 s3.1 s16 queries/min
Conversational query, sequential configuration88.1 s9.1 s50 queries/min

At eight-wide, the full 11,557-report corpus classifies in about 25 minutes. The complete Task A development pass over 10,114 reports cost about $4.20 in model inference; the whole evidence phase ran for under $30. Eight simultaneous users saw no latency degradation in the sequential configuration.

What this means for a buyer

  • No training data required to start. Zero-shot analysis matched a supervised model on this corpus. A new organization's reports can be coded on day one against its own vocabulary.

  • Every answer carries its evidence. The verbatim-span rate is measured, not asserted. A reviewer verifies against a quotation rather than approving an opaque output.

  • Reproducible. The corpus is public, the protocol is frozen, and the baselines are stated. An evaluator can rerun it.

  • Deployable where the data lives. The model interface is interchangeable, so the same pipeline runs against the model service an authorization boundary allows, including Government-authorized cloud regions, without a redesign.

Limits we state

ASRS is civil aviation safety reporting with NASA's taxonomy and sanitized single-narrative records. What transfers to other domains is the task shape: free-text operator reporting, expert-coded ground truth, and non-specialist users who need answers they can check. Text ingestion was tested; audio and graphics extraction were not. Held-out Task B used a smaller corpus than development, which makes set retrieval easier, and we report both sizes for that reason.

IP Network Solutions, Inc. · Leesburg, Virginia · Federal IT services and AI software since 2003.

Evaluation designed and performed by Michael Young, Program Director, AI Portfolio. Questions and demo requests: MYoung@ipnsinc.com