Notes
You pay for the whole conversation, every time
Why a long chat gets slower and more expensive with every turn, and what the API is actually doing when you hit send.
There is a mechanism underneath AI billing that surprises people the first time they see it on an invoice, and it is not complicated. It is just rarely spelled out.
A chat API has no memory. Each request is processed on its own. When you send your tenth message in a thread, the client does not send one message — it sends all ten, plus every reply, as one block of input. The model reads the whole conversation from the beginning, every single time.
That is what statelessness means in practice, and both Anthropic's and OpenAI's API documentation say so plainly: conversation history is resent with each request, and it is billed as input tokens on each request. Nothing is being stored on your behalf and recalled cheaply. It is being re-read.
What that does to cost
The thing to notice is the shape of the growth, not the unit price.
A thread where every turn is roughly the same size does not cost you a fixed amount per turn. Turn one pays for turn one. Turn two pays for turns one and two. Turn ten pays for all ten. The cost of a conversation grows with the square of its length, not in a straight line, because each new turn carries every turn before it.
This is why a long, meandering session can cost many times what the same work would have cost split across three focused threads — and why the fix is structural rather than a matter of being terser. You are not paying for what you typed. You are paying for everything still in the thread when you typed it.
One honest qualification. Prompt caching changes the arithmetic, and the major providers all offer some form of it: a prefix that has already been processed can be re-read at a reduced rate for a limited window. Where it applies and you have set it up, the growth is gentler than the plain version above. Two things stop that being a get-out, though. It discounts the re-reading; it does not stop it. And it rewards a stable prefix, so a thread that keeps changing its early turns gets less benefit than the headline rate suggests. Treat caching as a reason to structure a conversation deliberately, not as a reason to stop noticing its length.
What it does to latency
The same mechanism explains something people usually blame on the network.
Time-to-first-token is mostly prompt processing: the model has to read the input before it can begin producing output. A long conversation is a long input, so the pause before the first word appears grows as the thread grows. The tokens that follow arrive at roughly the same rate as they always did, which is why a long session feels like it stalls at the start and then recovers.
If a thread that was snappy this morning now takes several seconds to begin replying, and nothing about your connection changed, the input length is the first thing to look at.
What to do about it
Three habits, in order of how much they save.
Start a new thread when the task changes. This is the cheapest discipline in AI-assisted work and the one most people break by lunchtime. A fresh thread is not a loss of context; it is a deliberate choice about which context is worth carrying.
Carry a handover, not a transcript. When you do start fresh, bring the task, the constraints, the file paths and the decisions already made. Leave the argument you just had about naming. A six-line summary replaces forty turns and is usually more useful to the model than the forty turns were, because it has already had the noise removed.
Summarise rather than paste. A four-hundred-line log is not context, it is raw material. Three lines saying what failed, where, and what you have already ruled out is context. The log costs you twice — once to send and once in the attention it takes away from the thing you actually asked.
The part worth remembering
None of this is a reason to be frightened of long conversations. It is a reason to know what one costs, so that the length is a choice rather than an accident.
A thread is not a folder you are filing things in. It is a document that gets read aloud from the top, every time you add a line to the bottom.
Where to go next
- Context Management — what accumulates in a thread, and the five-question test for what belongs in one
- Multi-Turn Conversations — structuring a long task so it does not become a long thread
- Performance Tuning — the same arithmetic when the hardware is yours, including why the KV cache is allocated before you type anything