Context window
A context window is the maximum amount of text, measured in tokens, that a language model can take into account in a single request, including your instructions, any documents or conversation history, and the answer it writes. Anything outside the window is invisible to the model for that request, as if it never existed.
01why it matters for a business
Context windows have grown large enough that many models can take in hundreds of pages at once, which tempts teams to paste everything in and let the model sort it out. That works for some tasks but has three costs. Every token in the window is billed on every request. Long inputs slow responses. And models do not use all of a long input equally well; research has shown that details buried in the middle of a long prompt can be missed more often than details near the start or end.
So the window is a budget, not a storage system. Models do not remember past conversations unless your software feeds the history back in, and for large knowledge bases the efficient pattern is to retrieve the few relevant passages for each question rather than sending the whole library every time.
02what it looks like in practice
A legal operations team wants answers from a library of several thousand contracts. No context window holds all of them, and sending even a fraction per question would be slow and expensive. Instead the system indexes the library, retrieves the handful of clauses relevant to each question, and sends only those, with citations, to the model. For a single long agreement under negotiation, the team does the opposite: the whole document fits in the window, so the model reads it end to end.
03common mistakes
- Assuming the model remembers yesterday's conversation. Memory is something your software provides.
- Filling the window because you can. More context does not always mean better answers, and it always means more cost.
- Ignoring where key information sits in long prompts. Put instructions and the most important material where the model handles it best, then test.
- Not planning for documents that exceed the window.
04related terms
- TokenA token is the unit of text a language model reads and writes: a whole word, part of a word, a number, or a punctuation mark.
- Retrieval-augmented generation (RAG)Retrieval-augmented generation (RAG) is a technique in which an AI system first searches your own documents or data for passages relevant to a question, then gives those passages to a language model to write its answer.
- Inference costInference cost is what it costs to run an AI model to produce outputs, as opposed to the cost of training it.
- Large language model (LLM)A large language model (LLM) is an AI model trained on very large amounts of text to predict the next piece of text, which lets it write, summarize, translate, classify, extract information, and reason through problems in everyday language.
05where insomnia club fits
Insomnia Club designs how context is assembled for each request, deciding what to retrieve, what to summarize, and what to leave out, so systems stay accurate and affordable as your data grows.
see custom AI development →tell us what keeps you up at night.
Scoped by the people who ship it. Priced before we start.
book a call drop your number