I spent a solid afternoon hunting down the right quantized weights, ran my air-gap verification script clean, and got a 14B model loaded into VRAM on a 16GB card. Baseline sat at a comfortable 10GB, hello-world prompts came back instantly, and for about twenty minutes I thought the hard part was over.
Then I fed it an actual codebase instead of a toy prompt, let a multi-turn RAG loop run, and pushed the context out toward 32K tokens. The whole thing stuttered, tokens dropped to a crawl, and the process died with a CUDA out-of-memory error I hadn’t done anything obvious to deserve. I hadn’t changed the model. I hadn’t launched anything else. I’d just fallen into the trap every local AI setup eventually finds — assuming the model’s file size is the only number that matters.
Base Model Weights vs. The Expanding KV Cache #
Most people size their hardware off one number: the model’s static weight footprint. For a 4-bit quantized 8B model, that’s roughly params × 0.5 bytes — 8 billion times half a byte, about 4GB. That’s a fine starting estimate, though real quantization formats carry overhead the naive math misses; an actual Q4_K_M GGUF file runs closer to 0.6 bytes per parameter than a flat 0.5, because block-wise scale factors aren’t free. Close enough to plan around, not close enough to trust blindly.
Static weights are only half the equation, and they’re the easy half. The moment the model actually generates anything, it has to maintain a Key-Value cache — every token’s computed Key and Value vectors, stored across every attention layer and head, so the model never has to recompute the entire sequence from scratch just to produce the next token.
Unlike the weights, KV cache size isn’t fixed. It grows linearly with context length, batch size, and sequence depth, and the formula for it is straightforward once someone actually writes it down:
KV cache bytes = 2 × num_layers × num_kv_heads × head_dim × precision_bytes × context_length × batch_size`
The 2 accounts for storing both Key and Value. Plug in Llama 3 8B’s real published architecture — 32 layers, 8 KV heads under grouped-query attention, head dimension 128 — at FP16 precision, and each token costs exactly 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes, or 128KB. That number looks small right up until you multiply it by a context length with five digits.
The Hidden Overhead: Quantization, Dequantization, and CUDA Contexts #
Saved memory from quantization creates a false sense of security, because none of it accounts for what the runtime itself needs just to exist.
Spinning up a CUDA context — before a single tensor loads — typically claims somewhere around 500MB to 1.5GB of VRAM on its own, just for the runtime, cuDNN/cuBLAS handles, and driver bookkeeping. Tensor Cores also don’t do 4-bit math natively; a 4-bit model’s weights usually get unpacked and dequantized on the fly into FP16 or BF16 buffers before the actual matrix multiplication happens, which means the “4-bit model” is briefly a 16-bit model in memory during every forward pass. Add activation memory on top of that — intermediate tensors from the forward pass, scaling with batch size and how dense the context actually is — and the gap between “model file size” and “what VRAM actually needs” widens fast.
VRAM Saturation Across Context Lengths #
Here’s what that KV cache formula actually does to total VRAM demand as context grows, using an 8B model quantized to 4-bit (4GB base) with Llama 3 8B’s real architecture numbers:
| Context Length | Base Weights (4-bit) | KV Cache (FP16) | Total VRAM Needed |
|---|---|---|---|
| 2,000 tokens | 4.0 GB | ~0.26 GB | ~4.3 GB |
| 8,000 tokens | 4.0 GB | ~1.05 GB | ~5.1 GB |
| 32,000 tokens | 4.0 GB | ~4.19 GB | ~8.2 GB |
| 128,000 tokens | 4.0 GB | ~16.8 GB | ~20.8 GB |
Look at the bottom row. At 128K context, the KV cache alone runs more than four times the size of the entire quantized model. That’s not a rounding error creeping up — it’s the KV cache becoming the dominant memory consumer in the whole system, which is exactly the scenario that turns a comfortable 16GB card into an OOM crash the second a long agentic loop or a real codebase gets fed in.
Hardening Memory Management: Actionable Mitigations #
FlashAttention-2 / FlashInfer — Standard attention implementations materialize the full attention matrix, which scales quadratically with sequence length. FlashAttention tiles the computation instead, so the intermediate memory footprint for the attention step itself drops from O(N²) to O(N). It doesn’t shrink the KV cache — that’s still the formula above — but it stops the attention computation itself from being the thing that OOMs you first.
Quantized KV Caching (FP8 / INT4) — The KV cache defaults to FP16 unless told otherwise. Drop it to FP8 and it’s halved — 2 bytes to 1 is exactly 50%, because that’s arithmetic, not a vendor claim. Drop to INT4 and it’s a 75% reduction on the same logic, a quarter of a byte instead of two. Engines like vLLM and llama.cpp support this directly, with a real but generally small hit to generation coherence.
PagedAttention — Traditional CUDA allocation needs contiguous memory blocks, which means reserving for the worst-case sequence length whether it gets used or not. The original vLLM research measured this waste directly: existing systems lose 60-80% of KV cache memory to exactly this kind of fragmentation. PagedAttention, which is what vLLM actually runs on, breaks the cache into fixed-size blocks that don’t need to sit next to each other in memory — the same research reports that gets waste down under 4%. Not eliminated, but close enough that it stops being the bottleneck.
Strict Context Window Boundaries — Set a hard ceiling — `num_ctx` in Ollama, `–max-model-len` in vLLM — matched to what the VRAM budget in the table above can actually survive, not to what the model theoretically supports. A model advertising a 128K context window doesn’t mean the card in front of you can afford one.
The model’s file size was never the whole budget. It was the entry fee. The KV cache, the CUDA context overhead, the dequantization buffers — that’s the part of the bill that shows up after you’ve already committed, and it scales with how the model actually gets used, not with how big the download was.
Do the math on context length before finding out the hard way, mid-run, with a codebase half-fed
into a context window the card never had room for.
