Explainer
KV cache size per token: one formula does not fit every model
The KV cache formula, computed from real configs for 12 model types: GQA, MLA, sliding window, Mamba and linear hybrids. Capacity, prefix hits, handoff.
LLM inference, KV cache, Disaggregated serving, GPU infrastructure, Attention
The key-value (KV) cache holds the keys and values that attention has already computed, so a model does not compute them again for every new token. Its size per token is usually given as one formula: 2 × layers × KV heads × head dimension × bytes per element. For Qwen3-8B in BF16 that is 2 × 36 × 8 × 128 × 2 = 147,456 bytes, or 144 KiB, for every token of every request.
That formula is correct for Qwen3-8B. For many models released since 2024 it is wrong, each in a different way, and the error carries into every number that depends on it: how many requests fit on a GPU, whether a prefix cache hits, and how long the handoff takes between prefill and decode workers. This essay computes the cache for twelve model types from their published configurations, and shows where the formula breaks.
Point along the cache to change the context length.How big is the KV cache per token?
Five numbers decide it. Attention stores a key and a value (the 2) for every layer, for every KV head, across the head's dimension, at the size of one element.
\begin{aligned}
\text{KV bytes per token} &= 2 \times \text{layers} \times \text{KV heads} \\
&\quad \times \text{head dimension} \times \text{bytes per element}
\end{aligned}For Qwen3-8B in BF16:
2 \times 36 \times 8 \times 128 \times 2 = 147{,}456 \text{ bytes} = 144 \text{ KiB}
Point at an edge to name its factor; the last stops change the precision.Precision is the factor an operator controls. An FP8 cache holds the same tokens in half the memory, 73,728 bytes per token for Qwen3-8B, and FP4 in a quarter. The other four come from the model.
The per-token number grows quickly into the numbers that matter.
Point at a stage to see the cache at that scale.Why a long request can outweigh the model
Qwen3-8B has about 8.2 billion parameters, or 15.3 GiB of weights in BF16. Its native context is 32,768 tokens, where one request's cache is 4.5 GiB. With YaRN scaling to 131,072 tokens, one request's cache is 18 GiB: larger than the model that produced it.
How many requests fit on one GPU
Capacity is what remains after the weights. An H100 80GB reports about 79.6 GiB. With vLLM's gpu_memory_utilization set to 0.9 (its default before v0.20; 0.92 since), the engine takes 90 percent of that: about 71.6 GiB. Other engines count differently: SGLang computes mem_fraction_static at startup from the GPU and its settings, and TensorRT-LLM's free_gpu_memory_fraction (default 0.9) is a share of the memory left after the weights. The weights take 15.3 GiB. The remainder, about 56 GiB, also holds activations and CUDA graph buffers, so at most about 50 requests of 8K tokens fit at once, not the 64 that 72 GiB would suggest.
Point down the column to see each part of GPU memory.Two more effects change the real number. Engines allocate the cache in fixed blocks of tokens, so the last block of every request is partly empty. And tensor parallelism splits KV heads across GPUs only while there are enough heads to split.
Point at a request to see its last block.
Point along the row to change the number of GPUs.Here is the formula as code. It reads a model's config.json from Hugging Face and prints the first answer. The sections that follow show when that answer is wrong.
import json
import sys
config = json.load(open(sys.argv[1]))
config = config.get("text_config", config) # Multimodal configs nest the text model.
element_bytes = 2 # BF16. Use 1 for FP8.
if "kv_lora_rank" in config: # Latent attention (MLA): one shared row per token per layer.
per_layer = (config["kv_lora_rank"] + config["qk_rope_head_dim"]) * element_bytes
elif config.get("multi_query"): # Older configs, such as Falcon-7B, mark one shared KV head.
head_dim = config["hidden_size"] // config["num_attention_heads"]
per_layer = 2 * 1 * head_dim * element_bytes
else:
heads = config.get("num_key_value_heads", config["num_attention_heads"])
head_dim = config.get("head_dim") or config["hidden_size"] // config["num_attention_heads"]
per_layer = 2 * heads * head_dim * element_bytes
print(per_layer * config["num_hidden_layers"], "bytes per token, if every layer keeps a full cache")Which model types break the formula
Fewer KV heads: MHA, GQA and MQA
Older models gave every query head its own key and value head: multi-head attention (MHA). Grouped-query attention (GQA) lets several query heads share one KV head. Multi-query attention (MQA) lets all of them share a single one. The formula still holds; only the head count changes.
Point at a model to see how its query heads share KV heads.Llama-2-7B, with 32 KV heads, stores 524,288 bytes per token: 3.6 times Qwen3-8B. Concurrency estimates carried over between these designs are off by that factor.
Why MLA does not split across GPUs
DeepSeek-V3 uses multi-head latent attention (MLA). It stores one compressed row per token per layer: a 512-wide latent and a 64-wide position part. In decode, every query head reads that same row directly. Across 61 layers in BF16 that is 70,272 bytes per token, less than half of Qwen3-8B for a model about 80 times larger.
Point at the row, then at the heads that read it.The row belongs to no single head, so tensor parallelism cannot split it by head. Under tensor parallelism alone, each GPU holds the whole row, and per-GPU cache does not shrink as GPUs are added. Decode context parallelism (vLLM's --decode-context-parallel-size) shards the cache across tokens instead.
Point at the GPUs to add copies.DeepSeek-V3.2 adds sparse attention: a small indexer chooses which earlier tokens each new token reads. That reduces the work per step, not the storage. Every token is still stored, plus an indexer key, for 48,068 bytes per token in vLLM's FP8 layout for sparse MLA (fp8_ds_mla): the latent in FP8, the position part kept in BF16, and four scale values per row. Plain MLA, such as DeepSeek-V3, stores a standard FP8 row instead: 576 bytes per layer. Another engine that packs the FP8 row differently stores a slightly different number.
Point at a bar to compare what is stored with what is read.Sliding windows
In gpt-oss-120b, half of the 36 layers attend only to the last 128 tokens. Those layers stop growing once the window is full, at 4.5 MiB per request in total. The other 18 layers grow at 36,864 bytes per token. The formula, which counts every layer as full, overstates the cache by about two times at long context.
Point at a stack to compare the true cache with the counted one.Mixture of experts
A mixture-of-experts (MoE) model routes each token through a few of many feed-forward experts. The experts hold most of the weights, but they keep no cache. The cache is set by attention alone. Qwen3-30B-A3B, a model nearly four times larger than Qwen3-8B, stores 98,304 bytes per token, less than Qwen3-8B.
Point at the attention column, then at the experts.Hybrids keep a fixed state
Hybrid models replace most attention layers with layers that keep a fixed-size state instead of a growing cache: Mamba layers in Nemotron-H, linear attention in Qwen3-Next. RWKV-6 keeps no attention cache at all. What a request holds is no longer one thing: a small cache that grows, plus a state that does not.
Point at a model to see its state and its cache.Of Nemotron-H-8B's 52 layers, 4 are attention layers and 24 are Mamba layers; the rest hold nothing per request. Its cache grows at 16,384 bytes per token, and its state is 49.4 MiB per request with the state in BF16. With the state in BF16, it equals 3,162 tokens of its own cache, so for short requests the state is most of its memory, and at 8K tokens the request holds 177.4 MiB, still 6.5 times less than Qwen3-8B. vLLM keeps Nemotron-H's SSM state in FP32 by default, for accuracy: then the state is 97.4 MiB, it equals 6,234 tokens of cache, and an 8K request holds 225.4 MiB, 5.1 times less than Qwen3-8B.
Point along the line to change the context length.Sharing and images
Cross-layer attention (CLA) lets adjacent layers read one cache: sharing between pairs halves it, and the paper also tests groups of three. YOCO goes further: half of its layers read a single shared cache, and the other half keep a cache of constant size, a short window in its sliding-window variant. Vision-language models turn an image into ordinary cache positions: a 1024 by 1024 pixel image becomes 1,024 tokens, 144 MiB on Qwen3-VL-8B.
Point at a stack to compare separate and shared caches.
Point at the image, its tokens, and their cache.All twelve, side by side
| Model | Type | Bytes per token | Fixed state per request |
|---|---|---|---|
| Llama-2-7B | MHA | 524,288 | none |
| Qwen3-8B | GQA | 147,456 | none |
| Qwen3-VL-8B | GQA, images as tokens | 147,456 | none |
| Qwen3-30B-A3B | GQA, mixture of experts | 98,304 | none |
| DeepSeek-V3 | MLA | 70,272 | none |
| DeepSeek-V3.2 | MLA, sparse (vLLM FP8 layout) | 48,068 | none |
| gpt-oss-120b | GQA, sliding window | 36,864 | 4.5 MiB once the window is full |
| Qwen3-Next-80B-A3B | Linear attention hybrid | 24,576 | 37.7 MiB |
| Nemotron-H-8B | Mamba hybrid | 16,384 | 49.4 MiB in BF16 (97.4 MiB in vLLM's default FP32) |
| Falcon-7B | MQA | 8,192 | none |
| CLA models | Shared across layers | 2 times less (pairs), up to 3 | none |
| RWKV-6 | Recurrent | 0 | about 33 MiB for the 7B model (FP32 state) |
Point at a bar to name its model.Every number here is computed from public configurations, not measured. The BF16 bytes per token are model math, the same on every engine; DeepSeek-V3.2's FP8 row follows vLLM's sparse-MLA layout, and Nemotron-H's state is shown in BF16, where vLLM's default keeps it in FP32. Real engines add block padding, a memory reserve, CUDA graph buffers and, for FP8, scale tensors. Treat the table as the floor a measurement starts from. The data is available as CSV.
Why a hybrid model's prefix match is fragile
A prefix cache lets a new request reuse the cache of an earlier request that started with the same tokens, such as a long system prompt. The shared part is not computed again.
Point at the two requests to find their shared blocks.A hybrid model cannot reuse its attention cache alone. It also needs the fixed state as it was at the end of the shared part. vLLM saves that state only at a few block boundaries, and its blocks for hybrid models are large: 528 tokens for Qwen3.5-4B in BF16. In a plain single-node vLLM, if the saved state is missing, the attention match is thrown away too; a KV connector that accepts partial hybrid hits can keep it.
Point at the blocks, then at the saved state.vLLM issue #45238 shows how sharp the geometry is. With Qwen3.5-4B, 528-token blocks and prompts of about 2,200 tokens, a June 2026 build kept saved states at tokens 1,584 and 2,112. With a 1,600-token shared prefix, the state at 1,584 fell inside the shared part, and 52 of 64 requests hit. With a 1,500-token shared prefix, both fell in the unique part, and none hit. vLLM has since added a saved state at the detected end of the shared prefix (#37898), and the issue's author re-measured on a later build: the shorter shared prefix still hit only 25.5 percent of prefix tokens, against 54.0 percent for the longer one. The total loss is gone; half the reuse is still lost, and nothing in vLLM's metrics shows why. --prefix-match-unit can add saved states inside a short prefix, but it is off by default.
Point at each prompt to see where its saved state falls.Prefix caching on hybrids also has a cost when nothing hits. vLLM issue #60008 measured 13.4 to 16.2 percent lower output throughput (medians of six rounds; the issue's title says 11 to 16) on one NVFP4 Mamba hybrid on four GB200 GPUs, with prefix caching on and no hits at all, on a vLLM build with changes not yet merged. Later comments report a prototype that closes most of the gap. Both issues were open when this essay was published.
Point at each bar to compare throughput.Images add one more rule. Two requests with the same text and different images must not share a cache, so the cache key must include the image.
Point at the two images.What crosses the wire between prefill and decode
In disaggregated serving, one group of GPUs runs prefill and another runs decode, so the cache moves between them once per request. Its size is the per-token number times the prompt length, plus any fixed state.
Point at a model to send its cache.| Model | At 8K tokens | One 400G NIC | Eight 400G NICs | NVLink 4 |
|---|---|---|---|---|
| Llama-2-7B | 4.0 GiB | 86 ms | 10.7 ms | 9.5 ms |
| Qwen3-8B | 1.125 GiB | 24 ms | 3.0 ms | 2.7 ms |
| Qwen3-30B-A3B | 768 MiB | 16 ms | 2.0 ms | 1.8 ms |
| DeepSeek-V3 | 549 MiB | 11.5 ms | 1.4 ms | 1.3 ms |
| gpt-oss-120b | 292.5 MiB | 6.1 ms | 0.8 ms | 0.7 ms |
| Qwen3-Next-80B-A3B | 229.7 MiB | 4.8 ms | 0.6 ms | 0.5 ms |
| Nemotron-H-8B | 177.4 MiB | 3.7 ms | 0.5 ms | 0.4 ms |
Real transfers are slower than line rate.
Point at a link to send the same cache over it.The layout on each side matters as much as the size. When prefill and decode use different numbers of GPUs, KV heads must be regrouped in flight. For MLA, every decode GPU needs the whole row, so the same bytes are sent to each one.
Point at each side to see how the cache is laid out.When reloading beats recomputing
A cache that does not fit in GPU memory can move to host memory, a local SSD or a remote store, and come back when a request needs it.
Point at a tier to see what it holds.Moving it back is worth it only when it is faster than computing it again. Reload time is the cache size divided by the tier's bandwidth, plus a fixed cost to start the transfer. Recompute time is the prefill time for those tokens, which grows faster than the prompt, because attention work grows with the square of its length. So short prompts favor recomputing and long prompts favor reloading. A slower tier, or a larger cache per token, moves the crossover to longer prompts; an MLA or hybrid model, with a small cache per token, moves it to shorter ones. Where it falls depends on numbers that only your hardware gives.
Point at a prompt length to compare the two times.For the same reason, sparse attention does not make offload cheaper. It reads fewer tokens, but an offloaded cache must carry all of them.
Questions to ask any new model's config
Before you size a deployment for a new model, answer these from its config.json and its paper:
- What attention type does each layer use: full, sliding window, linear, Mamba or none?
- How many KV heads does each attention layer have, and is the head dimension stated separately?
- Does it store a latent row (MLA) instead of separate keys and values?
- Which layers have a window, and how long is it?
- Is there a fixed state per request, and how large is it in the dtype the engine uses?
- Do any layers share a cache?
- At your tensor-parallel size, is the cache split, or copied to every GPU?
The KV cache calculator runs these numbers for your model, context, GPUs and link.
Point at a model to see what a 128K-token request would hold.Questions engineers ask
Does an FP8 KV cache halve memory?
It halves the bytes per token, from 2 to 1 per element. Scale tensors add a small amount, and accuracy needs its own check for each model.
Why does my engine report fewer tokens than the formula?
The formula ignores block padding, the memory reserve, activations and CUDA graph buffers. Use the formula as the upper bound and the engine's startup log as the measurement.
Is the KV cache the same as the prefix cache?
No. The KV cache holds one request's keys and values. A prefix cache keeps those blocks after the request ends, so a later request with the same start can reuse them.
References
- Qwen3-8B config and the Qwen3 technical report
- Llama-2-7B config
- Falcon-7B config
- DeepSeek-V3 config; MLA in DeepSeek-V2
- DeepSeek-V3.2 config
- gpt-oss-120b config
- Qwen3-30B-A3B config
- Nemotron-H-8B config and the Nemotron-H paper
- Qwen3-Next-80B-A3B config
- Qwen3-VL-8B config
- RWKV-6: Eagle and Finch
- MQA: Fast Transformer Decoding; GQA: Ainslie et al.
- Mamba: Gu and Dao
- CLA: Brandon et al.; YOCO: Sun et al.
- vLLM issues #45238 and #60008, both open on 7 October 2026
- Related work: Sebastian Raschka's KV cache calculations, which covers the attention families; this essay adds precision, blocks, tensor-parallel copies, transfer cost and the prefix-cache behavior of hybrids.