highnet-rag

Sheet 2 ·How well does it work?

We searched for the answers to 105 questions whose answers we know exactly, through every search setting the pipeline offers. Every number on this page comes from that one run; none is typed in by hand.

Run
2026-09-28
Corpus build
5032bec50e35
Models called
rerank-2.5-lite, voyage-3.5-lite
Run cost
$0.12, paid separately from the visitor budget

Does the search find the answer?

For each of 105 answerable questions we know exactly where the answer sits in its article. recall@5 is the share of questions where a passage containing that answer was among the top 5 retrieved. The model can only answer from what it is given, so this caps everything that follows.

MRR (mean reciprocal rank) rewards finding the answer early: 1 when it is the top passage, ½ when second, ⅓ when third, and 0 when it is not in the top 10.

Measure
Small chunks
SearchReranker offReranker on
BM25recall@5 84%, over 105 questionsrecall@5 97%, over 105 questions
Vectorrecall@5 89%, over 105 questionsrecall@5 98%, over 105 questions
Hybridrecall@5 92%, over 105 questionsrecall@5 99%, over 105 questions
Medium chunks
SearchReranker offReranker on
BM25recall@5 93%, over 105 questionsrecall@5 99%, over 105 questions
Vectorrecall@5 90%, over 105 questionsrecall@5 99%, over 105 questions
Hybridrecall@5 96%, over 105 questionsrecall@5 99%, over 105 questions
Large chunks
SearchReranker offReranker on
BM25recall@5 94%, over 105 questionsrecall@5 99%, over 105 questions
Vectorrecall@5 91%, over 105 questionsrecall@5 98%, over 105 questions
Hybridrecall@5 97%, over 105 questionsrecall@5 99%, over 105 questions

Does the answer stick to the passages?

A judge model reads the question, the passages sent, and the answer sentence by sentence. It marks each sentence as supported by the passages or not, and the answer as correct or not against the known answer. A second, stronger model grades the first; it can still be wrong, so treat these as estimates.

Not measured in this run

Measuring this means generating an answer for every question and having a judge model grade it, which costs model calls. To keep eval spend to cents, this run measured retrieval only.

Does the agent help with two-part questions?

10 questions each need a fact from two different articles. One search tends to find one article; in agentic mode Claude splits the question and searches for each part.

Not measured in this run

Measuring this means generating an answer for every question and having a judge model grade it, which costs model calls. To keep eval spend to cents, this run measured retrieval only.

What the run cost

Where the questions come from

squad_auto
SQuAD 2.0 dev, picked automatically
150 · 105 answerable, 45 not
owner
Written by the owner
not written yet
compound
Cross-article, for agentic mode
10 · 10 answerable
Every call in the eval run, by model
ModelCallsTokens inTokens outCost
rerank-2.5-lite9455,793,1070$0.12
voyage-3.5-lite1051,2590< $0.0001
Total$0.12