Results
All 252 runs from one sweep: nine architectures over 28 questions, same corpus, same judge, same day. Every figure here is redrawn from the JSON the evaluation wrote, so nothing on this page can drift from the committed reports.
Scored 2026-08-28 23:59 UTC
Every metric, every stage
Best value per row is highlighted. Note that no single stage owns the whole table.
| Metric | v1 | v2 | v3 | v4 | v5 | v6 | v7 | v8 | v9 |
|---|---|---|---|---|---|---|---|---|---|
| Headline | 0.776 | 0.779 | 0.745 | 0.779 | 0.739 | 0.806 | 0.786 | 0.876 | 0.817 |
| Context precision | 0.402 | 0.481 | 0.501 | 0.512 | 0.472 | 0.717 | 0.596 | 0.571 | 0.571 |
| Context recall | 0.568 | 0.685 | 0.648 | 0.663 | 0.729 | 0.845 | 0.802 | 0.910 | 0.758 |
| MRR | 0.620 | 0.768 | 0.763 | 0.793 | 0.798 | 0.838 | 0.786 | 0.815 | 0.857 |
| nDCG | 0.666 | 0.809 | 0.810 | 0.812 | 0.839 | 0.876 | 0.820 | 0.830 | 0.841 |
| Hit rate | 0.857 | 0.964 | 0.964 | 0.964 | 1.000 | 1.000 | 0.964 | 1.000 | 0.964 |
| Faithfulness | 0.886 | 0.907 | 0.914 | 0.914 | 0.861 | 0.954 | 0.861 | 0.875 | 0.929 |
| Answer relevance | 0.954 | 0.943 | 0.911 | 0.889 | 0.900 | 0.875 | 0.939 | 0.989 | 0.889 |
| Answer correctness | 0.668 | 0.664 | 0.607 | 0.675 | 0.629 | 0.707 | 0.704 | 0.846 | 0.736 |
| Attack resisted | 0.60 | 0.60 | 0.00 | 0.60 | 0.40 | 0.40 | 0.60 | 0.80 | 1.00 |
Score across the nine stages
Correctness and faithfulness diverge at v9: it gets safer and less accurate at once.
Shape of each stage
Six metrics at once. A stage that is strong everywhere would fill the hexagon.
By question type
Global questions are the standing weakness of every stage, including the last one.
Quality against cost
Outlined points are on the frontier: nothing else is both cheaper and better. The expensive stages have to earn their place here, not just on accuracy.
What each stage costs to run
Agentic and multi-agent stages pay for their accuracy in tokens and seconds.
How each stage fails
Not just how often, but in which way. The mix shifts as the architectures change.
Retrieval recall as k grows
How much of the gold evidence is in the top k. This is retrieval quality alone, before the generator gets involved.
Retrieval precision as k grows
Stages that retrieve aggregate nodes trade precision for coverage by design.
Can the judge be trusted?
Each answer is judged three times and the median taken; this is how much the three passes disagreed. Where the spread approaches the gate's tolerance, differences between stages of that size are noise.
Every run
Nine stages by every question, coloured by headline score. Click a cell to read the answer, the judge's reasoning and what the stage did internally.
a01 | a02 | a03 | a04 | a05 | f01 | f02 | f03 | f04 | f05 | f06 | f07 | f08 | f09 | f10 | g01 | g02 | g03 | g04 | g05 | m01 | m02 | m03 | m04 | m05 | m06 | m07 | m08 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| v1 | ||||||||||||||||||||||||||||
| v2 | ||||||||||||||||||||||||||||
| v3 | ||||||||||||||||||||||||||||
| v4 | ||||||||||||||||||||||||||||
| v5 | ||||||||||||||||||||||||||||
| v6 | ||||||||||||||||||||||||||||
| v7 | ||||||||||||||||||||||||||||
| v8 | ||||||||||||||||||||||||||||
| v9 |
Per-tag leaders
Which architecture you would actually pick, per question type.
Router, injection guardrail, grader, critic and synthesiser as separate agents