Cover for the post “The five-question context test”

Packs

The five-question context test

Five questions to run over anything before it goes into a prompt. Free, no sign-up.

Run these over every piece of context before it goes into a prompt. They take about two seconds each once you have them in your head, and they are the whole method.

Ask Then
Does the model need this to produce a correct answer? Include it
Would it make a mistake without this? Include it
Nice to know, but not required? Leave it out
About a different part of the system? Leave it out
Could I summarise it in one sentence? Summarise it

Question five is the one that pays for itself. A log, a stack trace or a long prior discussion almost always compresses to a line or two — and the compressed version is better context than the original, because you have already decided what mattered.


The same request, twice

Too much:

Here's my entire project structure, all 47 files, the full package.json,
the complete README, three config files and last month's git log.

Now add a loading spinner to the submit button.

Enough:

In src/components/SubmitButton.tsx, line 23, handleSubmit needs a loading
state. We use the Spinner component from src/components/ui/Spinner.tsx.

Add an isLoading state that shows Spinner and disables the button during
submission.

The second is shorter, cheaper, faster — and produces better code.


Summarise, don't dump

Three patterns for the things people paste whole.

A failing test or stack trace

<test name> fails with <the actual error, one line>.
It happens in <file:line>.
I have already ruled out: <what you checked>.

A long log

Three lines from <log>: what failed, where, and what preceded it.
<paste only those three lines>

A prior conversation

Earlier we established: <decision 1>, <decision 2>.
Ruled out: <thing>, because <one-line reason>.
Current state: <what works, what doesn't>.

In all three, the model gets the same information and none of the noise. You are not withholding context. You are doing the triage yourself instead of asking it to.


Why volume costs you

Not an opinion — Chroma Research tested 18 models across varied input lengths and needle positions and found that performance degrades as input grows, "often in surprising and non-uniform ways", that even a single distractor reduces performance against baseline, and that focused prompts beat full ones in every model family tested.

The window is the most the model will accept. It was never a claim about where it does its best work.

Read the method rather than taking anyone's summary of it, including this one.


If you run models locally

Right-size the context window to the prompts you actually send. The KV cache is allocated up front from your configured length, not grown as you use it, so a 128K window on an 8B model can reserve roughly three times the memory the 4-bit weights take — about 16 GiB against about 4.9 GB — for no benefit if your prompts are 2K. Running the cache at q8_0 roughly halves it.


Where to go next