Sheet 2 ·How well does it work?
We searched for the answers to 105 questions whose answers we know exactly, through every search setting the pipeline offers. Every number on this page comes from that one run; none is typed in by hand.
- Run
- 2026-09-28
- Corpus build
- 5032bec50e35
- Models called
- rerank-2.5-lite, voyage-3.5-lite
- Run cost
- $0.12, paid separately from the visitor budget
Does the search find the answer?
For each of 105 answerable questions we know exactly where the answer sits in its article. recall@5 is the share of questions where a passage containing that answer was among the top 5 retrieved. The model can only answer from what it is given, so this caps everything that follows.
MRR (mean reciprocal rank) rewards finding the answer early: 1 when it is the top passage, ½ when second, ⅓ when third, and 0 when it is not in the top 10.
| Search | Reranker off | Reranker on |
|---|---|---|
| BM25 | recall@5 84%, over 105 questions | recall@5 97%, over 105 questions |
| Vector | recall@5 89%, over 105 questions | recall@5 98%, over 105 questions |
| Hybrid | recall@5 92%, over 105 questions | recall@5 99%, over 105 questions |
| Search | Reranker off | Reranker on |
|---|---|---|
| BM25 | recall@5 93%, over 105 questions | recall@5 99%, over 105 questions |
| Vector | recall@5 90%, over 105 questions | recall@5 99%, over 105 questions |
| Hybrid | recall@5 96%, over 105 questions | recall@5 99%, over 105 questions |
| Search | Reranker off | Reranker on |
|---|---|---|
| BM25 | recall@5 94%, over 105 questions | recall@5 99%, over 105 questions |
| Vector | recall@5 91%, over 105 questions | recall@5 98%, over 105 questions |
| Hybrid | recall@5 97%, over 105 questions | recall@5 99%, over 105 questions |
Does the answer stick to the passages?
A judge model reads the question, the passages sent, and the answer sentence by sentence. It marks each sentence as supported by the passages or not, and the answer as correct or not against the known answer. A second, stronger model grades the first; it can still be wrong, so treat these as estimates.
Not measured in this run
Measuring this means generating an answer for every question and having a judge model grade it, which costs model calls. To keep eval spend to cents, this run measured retrieval only.
Does the agent help with two-part questions?
10 questions each need a fact from two different articles. One search tends to find one article; in agentic mode Claude splits the question and searches for each part.
Not measured in this run
Measuring this means generating an answer for every question and having a judge model grade it, which costs model calls. To keep eval spend to cents, this run measured retrieval only.
What the run cost
Where the questions come from
- squad_auto
- SQuAD 2.0 dev, picked automatically
150 · 105 answerable, 45 not - owner
- Written by the owner
not written yet - compound
- Cross-article, for agentic mode
10 · 10 answerable
| Model | Calls | Tokens in | Tokens out | Cost |
|---|---|---|---|---|
| rerank-2.5-lite | 945 | 5,793,107 | 0 | $0.12 |
| voyage-3.5-lite | 105 | 1,259 | 0 | < $0.0001 |
| Total | $0.12 | |||