Skip to content

query expansion

Query expansion is a technique that reformulates or enriches a search query with extra terms before retrieval, so a system finds relevant documents even when the original wording misses them. A long-standing idea in information retrieval, it is now central to retrieval-augmented generation (RAG), where the retrieved context sets a ceiling on how good the model’s answer can be.

Expansion sits at the front of the retrieval pipeline that produces that context:

A raw query is expanded by classic and LLM methods, retrieved from a vector database, then re-ranked into the model's context.
Broaden for Recall, Re-Rank for Precision

Classic approaches add synonyms, stem words to their root forms, correct spelling, or reweight terms to sharpen the match. Pseudo-relevance feedback goes further, expanding a query with terms drawn from the top results of a first search pass.

In modern pipelines, a large language model often does the expanding. Multi-query generation rewrites one question into several phrasings, while HyDE, or hypothetical document embeddings, drafts a hypothetical answer and embeds that instead of the raw query to search a vector database.

The core trade-off is recall versus precision. Broadening a query surfaces more candidate documents but can also pull in noise, so retrieval systems often pair expansion with re-ranking to keep the strongest matches on top.

Python LlamaIndex: Step by Step RAG With Examples

Tutorial

LlamaIndex in Python: A RAG Guide With Examples

Learn how to set up LlamaIndex, choose an LLM, load your data, build and persist an index, and run queries to get grounded, reliable answers with examples.

intermediate ai

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 25, 2026