Transparency
Fleet RAG answer quality
Independent RAGAS evaluation with a disjoint GPT-5 judge (the answers are written by Claude Opus, so the judge is a different model family). Faithfulness is scored against the verbatim context the model actually retrieved. Each row is a RAG; the gauges show the mean score, with a tick at the pass threshold.
Pilot — 9 of 9 corpora evaluated so far. Rolling out across the fleet.
How each score is measured
Faithfulness
the AI's answer —checked vs— the retrieved sources
Every claim in the answer is grounded in a real retrieved source — no invented facts.
Answer relevancy
the AI's answer —checked vs— your question
The answer actually addresses what was asked, without padding or drift.
Context precision
the retrieved sources —checked vs— the reference answer
The right passages were retrieved and ranked first, not buried under noise.
Context recall
the reference answer —checked vs— the retrieved sources
All the evidence needed to answer was actually retrieved — nothing missed.
All runs: GPT-5 judge · Claude Opus answer generator. Each gauge fills to the score; the tick marks the pass threshold. Green = at/above, red = below.
For healthcare providers. AI-generated summaries may contain errors — verify against primary sources and clinical judgement. Not medical advice.