Reranking

Definition

Reranking is a second-stage scoring step that takes an initial candidate set retrieved cheaply (first stage) and rescores it using a more expensive, higher-quality model. The reranker sees both the query and document together, enabling richer interaction than first-stage retrievers.

Why Reranking Exists

First-stage retrievers (BM25, Bi-Encoder) must score millions of documents quickly — they use independent query/document representations or simple term statistics. This limits their expressiveness. A reranker only needs to score ~100–1000 candidates, so it can afford deep query-document interaction.

Standard Pipeline

Query
  │
  ▼
First-stage retrieval (BM25 / bi-encoder ANN)   → top-1000 candidates
  │
  ▼
Reranker (cross-encoder or LLM)                 → rescored top-1000
  │
  ▼
Final ranked list (top-10/20 served to user)

Reranker Types

Cross-Encoder (Most Common)

A Cross-Encoder processes query + document concatenated as a single sequence, producing a relevance score. Sees full interaction between query and document tokens.

  • Models: cross-encoder/ms-marco-MiniLM-L-6-v2, Cohere Rerank, bge-reranker-*
  • Latency: ~50–200ms for 100 candidates on GPU
  • Quality: significantly better than bi-encoder for nuanced relevance

LLM-as-Reranker

Use a large language model (LLM as Judge) to score or listwise-rank candidates. Higher quality, much higher cost.

  • Pointwise: “Is this document relevant to the query? Yes/No”
  • Listwise: “Rank these 10 documents by relevance”

ColBERT Late Interaction

ColBERT / Late Interaction sits between bi-encoder speed and cross-encoder quality — efficient enough for first-stage in some setups, but also used as a reranker.

Feature-Based (LTR) Re-ranker

A LTR model (LambdaMART) rescores the top-N using tabular features — BM25 components, popularity/CTR, price, freshness. Often deployed as an external secondary re-ranker outside the search engine, agnostic to retrieval. Metarank is the canonical open-source example; it trains on Implicit Judgments and serves with a ~20–30 ms latency budget. See Learn-to-Rank with OpenSearch and Metarank.

Structured-Decision Model

A general structured-output model can be used as a reranker without any ranking training: pose relevance as one typed true-or-false question per candidate against a shared state, and sort by the returned probability. Jev is the worked example — see Hev meets Jev, where this shape reaches 0.501 mean nDCG@10 across three BEIR subsets against 0.504 for Voyage rerank-3 and 0.404 for the unreranked BM25 order.

Two properties distinguish it from the families above. It returns a Calibrated Relevance Probability rather than a logit, so a single call can rerank and prune against a portable threshold. And the latency profile inverts: batching thirty documents into one call gives a good median but a long tail (p95 ~1.4 s vs under 0.5 s for hosted cross-encoders), while asking one question per pair fixes the tail at thirty times the requests.

Reranking in RAG

In RAG pipelines, reranking is critical: the LLM context window is limited, so only the top 3–5 chunks are included. A reranker narrows 50–100 retrieved chunks down to the best ones before LLM generation.

Tradeoffs

First-StageReranker
ThroughputMillions of docs/secHundreds of docs/sec
Latency<10ms50–500ms
QualityGoodExcellent
CostLowHigher
  • Retrieval Pipeline — the multi-stage architecture reranking fits into
  • Cross-Encoder — primary reranking architecture
  • Bi-Encoder — first-stage retriever that feeds the reranker
  • ColBERT — late interaction alternative
  • LLM as Judge — LLM-based reranking
  • RAG — key use case for reranking
  • Learning to Rank — related family of ranking approaches
  • MonoT5 — T5 pointwise neural reranker
  • RankLLaMA — LLaMA fine-tuned reranker
  • RankGPT — listwise LLM reranker
  • LambdaMART / Metarank — feature-based external secondary re-ranker
  • Feature Store — serves the features an LTR re-ranker consumes
  • Asymmetric Re-ranking — cheap second-phase recall recovery for binary-quantized retrieval, scoring a full-precision query against BQ document vectors
  • Kendall Rank Correlation — diagnostic for how much reranking changed the candidate order (not a quality metric)
  • Relational Transformer — reranks candidates by conditioning on structured relational data (typed fields, schema links) instead of text
  • Calibrated Relevance Probability — reranker output as a probability rather than an ordering-only score, which makes pruning and cross-leg comparison possible
  • Jev — a structured-decision model used as a reranker with no ranking training
  • hev-rerank — the minimal open-source implementation of that approach

When Reranking Becomes a System Boundary

From When Reranking Becomes a System Boundary (Ravindra Harige):

Retrieval defines eligibility; reranking defines order. If a document is not retrieved, no downstream stage can recover it.

Ranking as a Projection

Reranking does not redo retrieval-time computation (term matches, field contributions, BM25 components). It operates on a compressed, lossy representation of what survived into the candidate set. This is structural — not a failure of implementation.

Compensatory Reranking

The system crosses a boundary when performance gains come from widening the rerank window rather than improving retrieval. At that point the window size is load-bearing (not a latency knob), and reranking has become compensatory.

Evaluation Split

StageMetricBlindspot
RetrievalRecall@KNDCG can improve while recall is weak
RerankingNDCG, MRRMetrics improve while user-visible relevance plateaus

Retrieval (engineering) and reranking (ML/data science) are owned by different teams with different metrics. Neither dashboard shows the full picture. The gap closes only when someone is accountable for the space between them.

What Reranking Cannot Fix

A reranker re-scores the candidate set. It cannot reach outside it. If the correct document was never retrieved into the top-N, no reranker — and no downstream generator — recovers it; the document is simply gone, and the answer is produced from what remains.

This makes reranking a natural but frequently wrong place to debug. Attention gravitates there because that is where the visible output is, while the damage may already have been done at candidate selection. Before tuning a reranker, verify that the correct document is present in the candidate set at all rather than checking its rank — a distinct measurement, and the one that Recall@K is for.

A concrete instance where fusion arithmetic silently evicted the answer before reranking ran: Hybrid Fusion Failure - BM25 Displacing Reference Documents.

The Inverse Case: Perfect Recall, Ranking Still Fails

The argument above is about reranking being blamed for retrieval’s failures. The opposite case is worth holding alongside it, because it is diagnosed the same way and treated differently.

In TypeSafe Cookbook - Re-ranking, a BM25 top-30 shortlist over CLERC legal passages contained the correct passage for 100% of queries — retrieval was flawless — and the correct passage was still ranked first only 5% of the time before reranking, and 18% after. Recall@30 was not the constraint; nothing about widening the window would have helped.

So the same measurement that exonerates a reranker can also convict it. Checking whether the right document is in the candidate set tells you which stage owns the problem, and both answers are common:

Recall@KTop-1 accuracyWhere the problem is
LowLowRetrieval — reranking cannot reach what was never fetched
HighLowRanking — the candidate set is fine, the ordering judgment is hard
HighHighNeither

Legal citation matching lands in the second row because relevance there is propositional rather than topical: the gold passage must establish the specific rule the query invokes, and a passage on the same doctrine is a deliberate near-miss.

Topics