reranker
A reranker is a model that reorders an initial set of retrieved documents by how relevant each one is to a query, sharpening the results before they reach the application that will use them. It forms the second stage of a two-stage retrieval pipeline, most often within retrieval-augmented generation (RAG). In a RAG system, that application is a large language model, so the passages that survive the reordering are the ones that shape its answer.
The first stage uses a fast retriever, typically an embedding model over a vector database, to pull a broad candidate set with approximate nearest-neighbor search tuned for recall, favoring catching every relevant document over being selective. The reranker then scores each query-document pair for precision and sorts the candidates so the most relevant passages rise to the top, as this pipeline shows:
Most rerankers are cross-encoders: they pass the query and a document through a transformer together, so the two texts attend to each other and produce a single relevance score. This is more accurate than bi-encoder retrieval, which embeds each text separately, but far more expensive, because nothing can be precomputed.
That’s why a reranker runs only over the small candidate set from the first stage, not the whole document collection being searched. Hosted services such as Cohere Rerank and open models like the BGE reranker provide ready-made options.
Related Resources
Tutorial
Build an LLM RAG Chatbot With LangChain
Large language models (LLMs) have taken the world by storm, demonstrating unprecedented capabilities in natural language tasks. In this step-by-step tutorial, you'll leverage LLMs to build your own retrieval-augmented generation (RAG) chatbot using synthetic data with LangChain and Neo4j.
For additional information on related topics, take a look at the following resources:
- Embeddings and Vector Databases With ChromaDB (Tutorial)
- Vector Databases and Embeddings With ChromaDB (Course)
- LlamaIndex in Python: A RAG Guide With Examples (Tutorial)
- Using LlamaIndex for RAG in Python (Course)
- Hugging Face Transformers: Leverage Open-Source AI in Python (Tutorial)
- How to Use Ollama to Run Large Language Models Locally (Tutorial)
- First Steps With LangChain (Course)
- Build an LLM RAG Chatbot With LangChain (Quiz)
- Embeddings and Vector Databases With ChromaDB (Quiz)
- LlamaIndex in Python: A RAG Guide With Examples (Quiz)
- Hugging Face Transformers (Quiz)
- How to Use Ollama to Run Large Language Models Locally (Quiz)
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Martin Breuss • Updated Sept. 25, 2026