Rerankers Aren’t Magic Either: When the Cross-Encoder Layer Is Worth the Cost
Enterprise Document Intelligence [Vol. 1 #2bis] Why stacking a reranker on top of weak retrieval doesn’t save it, what cross-encoders actually fix vs what they don’t, and where the editorial position of the series lands.
by Rushikesh Gaikwad via Unsplash
This article tests the cost-perf gradient empirically: four embedding models from 2014 to 2024, plus three off-the-shelf cross-encoder rerankers, scored side by side on the cases Article 2 catalogued. The result is more surprising than the funnel suggests.
The seven models tested, with their license attestation URLs:
- GloVe-avg (2014, 300-dim word vectors): Apache 2.0, declared on the HuggingFace model card.
- all-MiniLM-L6-v2 (2021, 22M params, 384-dim): Apache 2.0, declared on the HuggingFace model card.
- text-embedding-ada-002 (OpenAI 2022, 1536-dim): proprietary; OpenAI Terms of Use.
- text-embedding-3-large (OpenAI 2024, 3072-dim): proprietary; OpenAI Terms of Use.
- bge-reranker-base (BAAI 2023, 278M params): MIT license, declared on the HuggingFace model card.
- bge-reranker-large (BAAI 2023, 560M params): MIT license, declared on the HuggingFace model card.
- cross-encoder/ms-marco-MiniLM-L-12-v2 (historical baseline): Apache 2.0, declared on the HuggingFace model card.
1. What a reranker actually is
1.1 The cost/precision gradient
- Bi-encoder embedding similarity. A precomputed vector per document. At query time the model encodes the query once and runs cosine similarity against the index. Milliseconds for millions of candidates. Cheap and approximate.
- Cross-encoder reranker. Query and passage are tokenised together and passed through a transformer that attends across both. The output is a single relevance score per pair. Cannot be precomputed because the query is part of the input. Tens of milliseconds per pair. Mid-cost, mid-precision.
- Chat-completion LLM. Reads a small candidate set and produces a structured answer. Hundreds of milliseconds, dollars per million tokens. Most expensive, most accurate.
1.2 The funnel
The architectural picture is a funnel. The corpus has, say, 200,000 pages. The embedding stage scores them all and returns the top 100. The reranker scores the 100 and returns the top 10. The LLM reads the 10 and produces an answer. Each arrow narrows the candidate pool by an order of magnitude or more.
1.3 Bi-encoder vs cross-encoder mechanically
A bi-encoder encodes the query and the passage independently. A cross-encoder tokenises query and passage together, allowing for fine-grained interactions.
2. The cost-perf gradient, tested on the same cases
...
2.1 Literal-token trap
Query hot dog, candidates: a food paraphrase (TARGET, zero shared tokens), the lexical trap the dog basked in the hot sun, and an unrelated decoy. 3-large is the only model that flips the trap to #2 and lifts the paraphrase to #1.
2.2 Synonym recovery
Query is green card needed. The right answer shares zero tokens with the query but is strict synonym, with the lexical trap sharing three. The grid shows an inversion of the cost-perf claim.
2.3 Topical proximity vs answer relevance
User question: “Who signed the contract?” The grid says that MiniLM is the only model that promotes the actual signature line to #1.
2.4 Signal dilution in long context
bge-large, bge-base, and ms-marco-MiniLM all rank the short answer #1 with the buried-answer paragraph #2.
2.5 The yes/no question
The one column that promotes the actual answer is ms-marco-MiniLM-L-12-v2.
3. Where the cross-encoder still breaks
Four failure modes that survive the cross-encoder layer regardless of size or family.
3.1 Negation
Does any cross-encoder pick up the inversion?
3.2 Exact identifiers
The intuition is that learned similarity will confuse identifiers.
3.3 Listing
A listing question wants all relevant items, not just the top-k.
3.4 Out-of-domain vocabulary
Specialised vocabularies often exist outside common distributions, leading to consistency issues.
4. Where rerankers actually justify their cost
The architectural choices that make rerankers mostly redundant (question parsing, classify-before-retrieve, expert keywords, specific pipelines for specific intents) are what the rest of the series builds.
5. Conclusion
The marginal dollar buys more lift at the embedding stage than the reranker stage on many query shapes.