Cover for the post “A 1M window is not an instruction to use 1M tokens”

Notes

A 1M window is not an instruction to use 1M tokens

Large context windows made it possible to paste everything. The research says relevance still beats volume — here is the method behind that claim.

The most common mistake in AI-assisted work is no longer giving the model too little. It is giving it everything.

Frontier context windows reached a million tokens, and a lot of people read that as permission. It is worth separating two different things the number could mean: the most the model will accept, and the amount at which it does its best work. The window is the first. It was never a claim about the second.

The two versions

Here is the same request, made twice.

The first: here is my entire project structure, all 47 files, the full package.json, the complete README, three config files and last month's git log. Now add a loading spinner to the submit button.

The second: in src/components/SubmitButton.tsx, line 23, handleSubmit needs a loading state. We use the Spinner component from src/components/ui/Spinner.tsx. Add an isLoading state that shows Spinner and disables the button during submission.

The second is shorter, cheaper, faster, and produces better code. That last part is the claim worth testing rather than asserting, so here is where it comes from.

What the research actually found

Chroma Research published Context Rot: How Increasing Input Tokens Impacts LLM Performance in July 2025. It is worth reading the method rather than the headline, because the method is what makes it useful.

They tested 18 models. The evaluations went beyond the standard needle-in-a-haystack format — extended variants across eight input lengths and eleven needle positions, tests that introduced distractors, and conversational question-answering at around 113,000 tokens.

Three findings bear directly on how you prompt:

  • Performance degrades as input length grows, and the paper is careful to say it does so "often in surprising and non-uniform ways". Not a cliff at the window limit — a decline on the way there.
  • "Even a single distractor reduces performance relative to the baseline." One irrelevant document. Not forty-six.
  • Focused prompts outperformed full ones across every model family tested.

That third finding is the one that matters most here, because it is the claim itself rather than a proxy for it. Dumping the codebase does not give the model more to work with. It gives it more to sift, and the one thing you cared about arrives diluted.

Read the paper if you are going to repeat any of this. It is linked at the bottom, and the method is the part people drop when they summarise it.

The test worth applying

Before any piece of context goes into a prompt, five questions. They take about two seconds each once you have them in your head.

  1. Does the model need this to produce a correct answer? Include it.
  2. Would it make a mistake without this? Include it.
  3. Nice to know, but not required? Leave it out.
  4. About a different part of the system? Leave it out.
  5. Could I summarise it in one sentence? Summarise it.

Question five is the one that pays for itself most often. A log, a stack trace or a long thread of prior discussion almost always compresses to a line or two, and the compressed version is better context than the original, because you have already done the work of deciding what mattered.

The hardware version of the same mistake

If you run models locally there is a second, more literal cost.

The KV cache is allocated up front from your configured context length, not grown as you use it. So a 128K window on an 8B model can reserve roughly three times the memory the quantised weights themselves take — for no benefit at all, if your prompts are 2K.

The arithmetic, since it is the sort of claim that deserves showing: an 8B model at 32 layers × 8 KV heads × 128 head dimensions × 2 (for K and V) × 2 bytes at fp16 works out at about 128 KiB per token, so about 16 GiB at 128K context, against roughly 4.9 GB for 4-bit weights. The multiple roughly halves if you run the cache at q8_0 precision, so the quantisation setting matters as much as the window size.

Right-size the window to the prompts you actually send.

The part worth remembering

The window is a ceiling. It was never a target.


Sources

Where to go next