← BACK TO EXPLORER
Transparency

Fleet RAG answer quality

Independent RAGAS evaluation with a disjoint GPT-5 judge (the answers are written by Claude Opus, so the judge is a different model family). Faithfulness is scored against the verbatim context the model actually retrieved. Each row is a RAG; the gauges show the mean score, with a tick at the pass threshold.

Pilot — 9 of 9 corpora evaluated so far. Rolling out across the fleet.

How each score is measured
Faithfulness
the AI's answer —checked vs— the retrieved sources
Every claim in the answer is grounded in a real retrieved source — no invented facts.
Answer relevancy
the AI's answer —checked vs— your question
The answer actually addresses what was asked, without padding or drift.
Context precision
the retrieved sources —checked vs— the reference answer
The right passages were retrieved and ranked first, not buried under noise.
Context recall
the reference answer —checked vs— the retrieved sources
All the evidence needed to answer was actually retrieved — nothing missed.
RAGFaithfulness
pass ≥ 0.75
Answer relevancy
pass ≥ 0.70
Context precision
pass ≥ 0.60
Context recall
pass ≥ 0.60
N
Extraintestinal (EIM)0.940.930.760.6710
IBD-PSC0.900.920.650.3710
IBD-Unclassified0.880.940.610.6310
Crohn's — Large Bowel0.870.900.750.6612
Pouchology0.850.880.650.7811
Crohn's — Ileocolic0.820.910.750.5610
Ulcerative Colitis0.800.870.620.6010
Crohn's — Small Bowel0.780.840.760.6911
Perianal Crohn's0.690.920.760.7511

All runs: GPT-5 judge · Claude Opus answer generator. Each gauge fills to the score; the tick marks the pass threshold. Green = at/above, red = below.

For healthcare providers. AI-generated summaries may contain errors — verify against primary sources and clinical judgement. Not medical advice.