I had a local agent loop running last week — nothing fancy, read a file, search a directory, format the output — and the first step came back in about half a second. By the fourth step, after the thing had read a config file and searched a folder, every single action was taking somewhere north of ten seconds before it emitted a single token. Same model, same hardware. Nothing had gotten slower except the number of times I’d asked it to do something.
The obvious assumption is that the model itself slowed down, and that’s not what happened. Token generation speed — the decode phase — never changed. What actually happened is that the agent framework was re-sending the entire, ever-growing conversation on every iteration, forcing the inference engine to reprocess thousands of tokens from scratch before it could even start generating the next response. That reprocessing step is called prefill, and it behaves nothing like decode, which is exactly why it’s the one that quietly ate the whole loop.
Prefill (Compute-Bound) vs. Decode (Memory-Bound) #
Prefill is time-to-first-token — the phase where the model computes the Key-Value cache for every incoming context token before generating anything. It’s highly parallelized matrix math; the GPU can chew through all the context tokens at once, and the cost scales directly with how many tokens are sitting in that context.
Decode is the opposite kind of problem. Once the KV cache exists, generating each new token is sequential by nature — token 51 can’t get computed before token 50 exists — and it’s bound by memory bandwidth, not compute, since every step means reading the full model weights and cache back off memory just to produce one token. That’s why decode speed stays roughly flat regardless of how long the conversation already is, while prefill gets more expensive every single time the context grows.
In a normal single-shot generation, prefill happens exactly once. In an agent loop, it happens once per step, over a context that keeps expanding, so the total prefill work across N steps grows a lot faster than N — and it’s the part of the system nobody’s watching until it quietly dominates everything else.
The Architecture Fixes #
Automatic prefix caching is the real fix, and it isn’t literally the same technique across every runtime even though people talk about it like one thing. SGLang’s version is specifically called RadixAttention — it tracks shared prompt prefixes in a radix tree and reuses the cached KV state for anything already seen. vLLM runs its own automatic prefix caching doing the same conceptual job through a different implementation, not branded RadixAttention. Either way, the point’s identical: identical prefixes — system instructions, tool definitions, everything from prior turns that hasn’t changed — stay cached in memory between steps, so the server only has to prefill the actual delta, the new tool result that just came back, instead of the entire history all over again.
Chunked prefill handles a different problem: one enormous prefill can block decode for other concurrent requests on a shared server, which is what causes visible stuttering. Splitting the prefill into smaller chunks — 512 or 2048 tokens at a time — and interleaving them with ongoing decode steps keeps the stream from stalling out while a big prompt gets processed.
Context trimming is the low-tech fix that still matters most days: don’t hand the model a thousand-line raw JSON blob or an entire HTML dump as “tool output” when a five-line summary would do the same job. Every unstructured byte left sitting in the context is something that has to get prefilled again on the next step, cache or no cache, the moment anything upstream of it changes.
What Caching Actually Saves — In Tokens, Not Guesses #
In a standard tool-calling loop, every subsequent turn appends new execution results to the context history. Without prefix caching, the inference engine must reprocess the entire context from token zero during every prefill phase. The table below illustrates the exact token recomputation workload across a 4-step agent sequence:
| Agent Step | Cumulative Context | Tokens Reprocessed (No Cache) | Tokens Reprocessed (With Prefix Cache) | Reduction |
|---|---|---|---|---|
| Step 1: Initial Plan | ~500 tokens | 500 | 500 (nothing cached yet) | — |
| Step 2: Tool Exec & Fetch | ~1,500 tokens | 1,500 | 1,000 (delta only) | 33% fewer tokens |
| Step 3: Reasoning & Parse | ~3,500 tokens | 3,500 | 2,000 (delta only) | 43% fewer tokens |
| Step 4: Final Synthesis | ~6,000 tokens | 6,000 | 2,500 (delta only) | 58% fewer tokens |
Without caching, cumulative tokens reprocessed across all four steps comes to 11,500. With caching, it’s 6,000 — which isn’t a coincidence, it’s exactly the final context length, because with perfect prefix caching every token only ever gets prefilled once, the moment it’s first added. Actual wall-clock seconds depend on real hardware throughput and need measuring on the actual machine running this, not asserting in an article. The token math doesn’t.
Verify It on Your Own Hardware #
Step Prompt tok Prefill (s) Decode (s)
1 X X.XX X.XX
2 X X.XX X.XX
3 X X.XX X.XX
4 X X.XX X.XX
Watch the Prefill column, not the Decode column. Prefill should climb step over step as prompt tokens pile up. Decode should stay roughly flat, since it’s generating a similar amount of output each time regardless of how much history came before it. If Prefill isn’t climbing on your setup, either the context genuinely isn’t growing the way it looks like it is, or something’s already caching more than expected.
Once the problem’s confirmed real on actual hardware, the fix is a flag, not a rewrite: `–enable-prefix-caching` in vLLM, RadixAttention on by default in SGLang. Check the server’s debug logs for cache hit rate after turning it on — that number climbing toward the size of the system prompt is the actual proof it’s working, not just a config file that says it should be.
The model was never the bottleneck. The conversation history was, and it got worse every step because nobody told the server it was allowed to remember what it already computed five seconds ago.
Cache the prefix, trim the garbage out of tool output before it hits the context, and check the logs instead of assuming the flag did what the documentation says it does.
