context window
A context window is the maximum span of tokens an LLM can jointly attend to in one pass. Everything in the request counts toward it: the system prompt, the conversation history, tool definitions and tool results, any attached files or images, and your prompt. The output the model generates counts too, including its internal reasoning tokens.
Because this window is finite, overly long inputs or requested responses may be truncated or refused, so systems often summarize, chunk, or retrieve relevant subsets to fit the limit. Vendors specify their maximum window size, and the major providers all count prompt plus response tokens—reasoning tokens included—toward that limit.
Some modern models support very large context windows, from hundreds of thousands up to millions of tokens, though in practice recall and accuracy degrade as the window fills, a phenomenon now called context rot.
Mitigating this degradation is mainly a context engineering problem: agent harnesses compact or summarize older turns, clear stale tool results, offload state to external notes or memory, and retrieve only the slices they need. Meanwhile, model-side work on positional encoding and attention keeps pushing the ceiling higher.
Related Resources
Tutorial
Context Engineering for Python Codebases
Learn how context engineering shapes what your AI coding agent sees on every turn, and use four practical strategies to keep your Python projects on track.
For additional information on related topics, take a look at the following resources:
- Build an LLM RAG Chatbot With LangChain (Tutorial)
- Context Engineering for Python Codebases (Quiz)
- First Steps With LangChain (Course)
- Build an LLM RAG Chatbot With LangChain (Quiz)
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Leodanis Pozo Ramos • Updated Sept. 21, 2026