Rerankers Aren’t Magic Either: When the Cross-Encoder Layer Is Worth the Cost

Enterprise Document Intelligence [Vol. 1 #2bis] Why stacking a reranker on top of weak retrieval doesn’t save it, what cross-encoders actually fix vs what they don’t, and where the editorial position of the series lands.

by Rushikesh Gaikwad via Unsplash

This article tests the cost-perf gradient empirically: four embedding models from 2014 to 2024, plus three off-the-shelf cross-encoder rerankers, scored side by side on the cases Article 2 catalogued. The result is more surprising than the funnel suggests.

The seven models tested, with their license attestation URLs:

1. What a reranker actually is

1.1 The cost/precision gradient

  1. Bi-encoder embedding similarity. A precomputed vector per document. At query time the model encodes the query once and runs cosine similarity against the index. Milliseconds for millions of candidates. Cheap and approximate.
  2. Cross-encoder reranker. Query and passage are tokenised together and passed through a transformer that attends across both. The output is a single relevance score per pair. Cannot be precomputed because the query is part of the input. Tens of milliseconds per pair. Mid-cost, mid-precision.
  3. Chat-completion LLM. Reads a small candidate set and produces a structured answer. Hundreds of milliseconds, dollars per million tokens. Most expensive, most accurate.

1.2 The funnel

The architectural picture is a funnel. The corpus has, say, 200,000 pages. The embedding stage scores them all and returns the top 100. The reranker scores the 100 and returns the top 10. The LLM reads the 10 and produces an answer. Each arrow narrows the candidate pool by an order of magnitude or more.

1.3 Bi-encoder vs cross-encoder mechanically

A bi-encoder encodes the query and the passage independently. A cross-encoder tokenises query and passage together, allowing for fine-grained interactions.

2. The cost-perf gradient, tested on the same cases

...

2.1 Literal-token trap

Query hot dog, candidates: a food paraphrase (TARGET, zero shared tokens), the lexical trap the dog basked in the hot sun, and an unrelated decoy. 3-large is the only model that flips the trap to #2 and lifts the paraphrase to #1.

2.2 Synonym recovery

Query is green card needed. The right answer shares zero tokens with the query but is strict synonym, with the lexical trap sharing three. The grid shows an inversion of the cost-perf claim.

2.3 Topical proximity vs answer relevance

User question: “Who signed the contract?” The grid says that MiniLM is the only model that promotes the actual signature line to #1.

2.4 Signal dilution in long context

bge-large, bge-base, and ms-marco-MiniLM all rank the short answer #1 with the buried-answer paragraph #2.

2.5 The yes/no question

The one column that promotes the actual answer is ms-marco-MiniLM-L-12-v2.

3. Where the cross-encoder still breaks

Four failure modes that survive the cross-encoder layer regardless of size or family.

3.1 Negation

Does any cross-encoder pick up the inversion?

3.2 Exact identifiers

The intuition is that learned similarity will confuse identifiers.

3.3 Listing

A listing question wants all relevant items, not just the top-k.

3.4 Out-of-domain vocabulary

Specialised vocabularies often exist outside common distributions, leading to consistency issues.

4. Where rerankers actually justify their cost

The architectural choices that make rerankers mostly redundant (question parsing, classify-before-retrieve, expert keywords, specific pipelines for specific intents) are what the rest of the series builds.

5. Conclusion

The marginal dollar buys more lift at the embedding stage than the reranker stage on many query shapes.

6. Further reading