Packs
The five-question context test
Five questions to run over anything before it goes into a prompt. Free, no sign-up.
Run these over every piece of context before it goes into a prompt. They take about two seconds each once you have them in your head, and they are the whole method.
| Ask | Then |
|---|---|
| Does the model need this to produce a correct answer? | Include it |
| Would it make a mistake without this? | Include it |
| Nice to know, but not required? | Leave it out |
| About a different part of the system? | Leave it out |
| Could I summarise it in one sentence? | Summarise it |
Question five is the one that pays for itself. A log, a stack trace or a long prior discussion almost always compresses to a line or two — and the compressed version is better context than the original, because you have already decided what mattered.
The same request, twice
Too much:
Here's my entire project structure, all 47 files, the full package.json,
the complete README, three config files and last month's git log.
Now add a loading spinner to the submit button.
Enough:
In src/components/SubmitButton.tsx, line 23, handleSubmit needs a loading
state. We use the Spinner component from src/components/ui/Spinner.tsx.
Add an isLoading state that shows Spinner and disables the button during
submission.
The second is shorter, cheaper, faster — and produces better code.
Summarise, don't dump
Three patterns for the things people paste whole.
A failing test or stack trace
<test name> fails with <the actual error, one line>.
It happens in <file:line>.
I have already ruled out: <what you checked>.
A long log
Three lines from <log>: what failed, where, and what preceded it.
<paste only those three lines>
A prior conversation
Earlier we established: <decision 1>, <decision 2>.
Ruled out: <thing>, because <one-line reason>.
Current state: <what works, what doesn't>.
In all three, the model gets the same information and none of the noise. You are not withholding context. You are doing the triage yourself instead of asking it to.
Why volume costs you
Not an opinion — Chroma Research tested 18 models across varied input lengths and needle positions and found that performance degrades as input grows, "often in surprising and non-uniform ways", that even a single distractor reduces performance against baseline, and that focused prompts beat full ones in every model family tested.
The window is the most the model will accept. It was never a claim about where it does its best work.
Read the method rather than taking anyone's summary of it, including this one.
If you run models locally
Right-size the context window to the prompts you actually send. The KV cache is
allocated up front from your configured length, not grown as you use it, so a
128K window on an 8B model can reserve roughly three times the memory the 4-bit
weights take — about 16 GiB against about 4.9 GB — for no benefit if your prompts
are 2K. Running the cache at q8_0 roughly halves it.
Where to go next
- A 1M window is not an instruction to use 1M tokens — the argument, with the research method
- Context Management — the full lesson
- Performance Tuning — prompt processing and
num_ctx