Debarshi Das

Explainer

KV cache size per token: one formula does not fit every model

The KV cache formula, computed from real configs for 12 model types: GQA, MLA, sliding window, Mamba and linear hybrids. Capacity, prefix hits, handoff.

LLM inference, KV cache, Disaggregated serving, GPU infrastructure, Attention

The key-value (KV) cache holds the keys and values that attention has already computed, so a model does not compute them again for every new token. Its size per token is usually given as one formula: 2 × layers × KV heads × head dimension × bytes per element. For Qwen3-8B in BF16 that is 2 × 36 × 8 × 128 × 2 = 147,456 bytes, or 144 KiB, for every token of every request.

That formula is correct for Qwen3-8B. For many models released since 2024 it is wrong, each in a different way, and the error carries into every number that depends on it: how many requests fit on a GPU, whether a prefix cache hits, and how long the handoff takes between prefill and decode workers. This essay computes the cache for twelve model types from their published configurations, and shows where the formula breaks.

Two blocks side by side: the model's weights, and a taller KV cache block for one long request.Point along the cache to change the context length.
Qwen3-8B: at 128K tokens (with YaRN), one request's cache outweighs the model itself.

How big is the KV cache per token?

Five numbers decide it. Attention stores a key and a value (the 2) for every layer, for every KV head, across the head's dimension, at the size of one element.

KV bytes per token=2×layers×KV heads×head dimension×bytes per element\begin{aligned} \text{KV bytes per token} &= 2 \times \text{layers} \times \text{KV heads} \\ &\quad \times \text{head dimension} \times \text{bytes per element} \end{aligned}

For Qwen3-8B in BF16:

2×36×8×128×2=147,456 bytes=144 KiB2 \times 36 \times 8 \times 128 \times 2 = 147{,}456 \text{ bytes} = 144 \text{ KiB}
A block measured by dimension lines, squashed to a quarter of its height under a dashed outline of its full size.Point at an edge to name its factor; the last stops change the precision.
Five numbers multiply to 147,456 bytes per token. Precision is the only one you can change freely.

Precision is the factor an operator controls. An FP8 cache holds the same tokens in half the memory, 73,728 bytes per token for Qwen3-8B, and FP4 in a quarter. The other four come from the model.

The per-token number grows quickly into the numbers that matter.

A stack of thin slabs, one tall column of them, and a grid of sixty-four such columns, left to right.Point at a stage to see the cache at that scale.
Each token adds 144 KiB. One 8K request holds 1.125 GiB; 64 of them hold 72 GiB.

Why a long request can outweigh the model

Qwen3-8B has about 8.2 billion parameters, or 15.3 GiB of weights in BF16. Its native context is 32,768 tokens, where one request's cache is 4.5 GiB. With YaRN scaling to 131,072 tokens, one request's cache is 18 GiB: larger than the model that produced it.

How many requests fit on one GPU

Capacity is what remains after the weights. An H100 80GB reports about 79.6 GiB. With vLLM's gpu_memory_utilization set to 0.9 (its default before v0.20; 0.92 since), the engine takes 90 percent of that: about 71.6 GiB. Other engines count differently: SGLang computes mem_fraction_static at startup from the GPU and its settings, and TensorRT-LLM's free_gpu_memory_fraction (default 0.9) is a share of the memory left after the weights. The weights take 15.3 GiB. The remainder, about 56 GiB, also holds activations and CUDA graph buffers, so at most about 50 requests of 8K tokens fit at once, not the 64 that 72 GiB would suggest.

A GPU's memory as a column: weights at the bottom, a pool of request slots, and dashed slots that do not fit above it.Point down the column to see each part of GPU memory.
On one 80 GB GPU, after the weights, at most about 50 such requests fit, not 64.

Two more effects change the real number. Engines allocate the cache in fixed blocks of tokens, so the last block of every request is partly empty. And tensor parallelism splits KV heads across GPUs only while there are enough heads to split.

Two rows of equal blocks of different lengths; the last block in each row is only partly filled.Point at a request to see its last block.
Memory is handed out in fixed blocks, so each request's last block is partly empty.
Sixteen GPUs in two rows, one slab on each; the second row's slabs are copies of the first.Point along the row to change the number of GPUs.
Eight KV heads split across up to eight GPUs. At sixteen, every head is held twice.

Here is the formula as code. It reads a model's config.json from Hugging Face and prints the first answer. The sections that follow show when that answer is wrong.

Python
import json
import sys

config = json.load(open(sys.argv[1]))
config = config.get("text_config", config)  # Multimodal configs nest the text model.
element_bytes = 2  # BF16. Use 1 for FP8.

if "kv_lora_rank" in config:  # Latent attention (MLA): one shared row per token per layer.
    per_layer = (config["kv_lora_rank"] + config["qk_rope_head_dim"]) * element_bytes
elif config.get("multi_query"):  # Older configs, such as Falcon-7B, mark one shared KV head.
    head_dim = config["hidden_size"] // config["num_attention_heads"]
    per_layer = 2 * 1 * head_dim * element_bytes
else:
    heads = config.get("num_key_value_heads", config["num_attention_heads"])
    head_dim = config.get("head_dim") or config["hidden_size"] // config["num_attention_heads"]
    per_layer = 2 * heads * head_dim * element_bytes

print(per_layer * config["num_hidden_layers"], "bytes per token, if every layer keeps a full cache")

Which model types break the formula

Fewer KV heads: MHA, GQA and MQA

Older models gave every query head its own key and value head: multi-head attention (MHA). Grouped-query attention (GQA) lets several query heads share one KV head. Multi-query attention (MQA) lets all of them share a single one. The formula still holds; only the head count changes.

A rail of query heads wired down to one of three rows of KV heads; the rows shrink from long to a sliver.Point at a model to see how its query heads share KV heads.
Fewer KV heads, smaller cache: 512 KiB, 144 KiB and 8 KiB per token.

Llama-2-7B, with 32 KV heads, stores 524,288 bytes per token: 3.6 times Qwen3-8B. Concurrency estimates carried over between these designs are off by that factor.

Why MLA does not split across GPUs

DeepSeek-V3 uses multi-head latent attention (MLA). It stores one compressed row per token per layer: a 512-wide latent and a 64-wide position part. In decode, every query head reads that same row directly. Across 61 layers in BF16 that is 70,272 bytes per token, less than half of Qwen3-8B for a model about 80 times larger.

A long blue bar cut near one end, a row of small heads behind it, and a dashed line from every head to the bar.Point at the row, then at the heads that read it.
MLA stores one shared row per token per layer, and every head reads it: 70,272 bytes per token.

The row belongs to no single head, so tensor parallelism cannot split it by head. Under tensor parallelism alone, each GPU holds the whole row, and per-GPU cache does not shrink as GPUs are added. Decode context parallelism (vLLM's --decode-context-parallel-size) shards the cache across tokens instead.

Eight GPUs in a row, each holding the same wide slab; one is the original and seven are copies.Point at the GPUs to add copies.
MLA's single row cannot be split by head, so under tensor parallelism eight GPUs hold eight copies.

DeepSeek-V3.2 adds sparse attention: a small indexer chooses which earlier tokens each new token reads. That reduces the work per step, not the storage. Every token is still stored, plus an indexer key, for 48,068 bytes per token in vLLM's FP8 layout for sparse MLA (fp8_ds_mla): the latent in FP8, the position part kept in BF16, and four scale values per row. Plain MLA, such as DeepSeek-V3, stores a standard FP8 row instead: 576 bytes per layer. Another engine that packs the FP8 row differently stores a slightly different number.

A tall stored bar in blue topped by an orange key segment, beside a short read bar with the same orange segment.Point at a bar to compare what is stored with what is read.
Sparse attention reads fewer tokens but still stores all of them, plus a small extra key: 48,068 bytes.

Sliding windows

In gpt-oss-120b, half of the 36 layers attend only to the last 128 tokens. Those layers stop growing once the window is full, at 4.5 MiB per request in total. The other 18 layers grow at 36,864 bytes per token. The formula, which counts every layer as full, overstates the cache by about two times at long context.

A stack of 36 thin layers with a blue bar beside every other one and a short stub beside the rest.Point at a stack to compare the true cache with the counted one.
gpt-oss-120b: half its layers keep only the last 128 tokens. Counting them as full overstates the cache.

Mixture of experts

A mixture-of-experts (MoE) model routes each token through a few of many feed-forward experts. The experts hold most of the weights, but they keep no cache. The cache is set by attention alone. Qwen3-30B-A3B, a model nearly four times larger than Qwen3-8B, stores 98,304 bytes per token, less than Qwen3-8B.

An attention tower wired to a short blue cache column, beside a wide tower of expert slabs with no wire.Point at the attention column, then at the experts.
Experts do not touch the cache: Qwen3-30B-A3B holds 96 KiB per token, less than Qwen3-8B.

Hybrids keep a fixed state

Hybrid models replace most attention layers with layers that keep a fixed-size state instead of a growing cache: Mamba layers in Nemotron-H, linear attention in Qwen3-Next. RWKV-6 keeps no attention cache at all. What a request holds is no longer one thing: a small cache that grows, plus a state that does not.

A layer tower with a few bright dots, a blue box of fixed state, and a short stack of cache slabs beside it.Point at a model to see its state and its cache.
Hybrids keep a fixed state per request beside a small growing cache; RWKV keeps only the state.

Of Nemotron-H-8B's 52 layers, 4 are attention layers and 24 are Mamba layers; the rest hold nothing per request. Its cache grows at 16,384 bytes per token, and its state is 49.4 MiB per request with the state in BF16. With the state in BF16, it equals 3,162 tokens of its own cache, so for short requests the state is most of its memory, and at 8K tokens the request holds 177.4 MiB, still 6.5 times less than Qwen3-8B. vLLM keeps Nemotron-H's SSM state in FP32 by default, for accuracy: then the state is 97.4 MiB, it equals 6,234 tokens of cache, and an 8K request holds 225.4 MiB, 5.1 times less than Qwen3-8B.

A flat ribbon and a rising wedge crossing at a small blue block, with a steep thin wedge behind leaving the top.Point along the line to change the context length.
With the state in BF16: below about 3,200 tokens the fixed state is most of its memory; at 8K it uses 6.5x less.

Sharing and images

Cross-layer attention (CLA) lets adjacent layers read one cache: sharing between pairs halves it, and the paper also tests groups of three. YOCO goes further: half of its layers read a single shared cache, and the other half keep a cache of constant size, a short window in its sliding-window variant. Vision-language models turn an image into ordinary cache positions: a 1024 by 1024 pixel image becomes 1,024 tokens, 144 MiB on Qwen3-VL-8B.

Six stacked layers with three blue cache blocks, each shared by a pair of layers through a dashed link.Point at a stack to compare separate and shared caches.
Cross-layer sharing lets adjacent layers read one cache: half the size when pairs share.
A framed square tile cut into a fine grid, its rows stacked into a blue wall at the back of the frame.Point at the image, its tokens, and their cache.
A 1024 by 1024 pixel image becomes 1,024 tokens: 144 MiB of cache.

All twelve, side by side

KV cache per token and fixed state per request, BF16 unless stated, computed from each model's config.json.

ModelType

Bytes per token

Fixed state per request
Llama-2-7BMHA524,288none
Qwen3-8BGQA147,456none
Qwen3-VL-8BGQA, images as tokens147,456none
Qwen3-30B-A3BGQA, mixture of experts98,304none
DeepSeek-V3MLA70,272none
DeepSeek-V3.2MLA, sparse (vLLM FP8 layout)48,068none
gpt-oss-120bGQA, sliding window36,8644.5 MiB once the window is full
Qwen3-Next-80B-A3BLinear attention hybrid24,57637.7 MiB
Nemotron-H-8BMamba hybrid16,38449.4 MiB in BF16 (97.4 MiB in vLLM's default FP32)
Falcon-7BMQA8,192none
CLA modelsShared across layers2 times less (pairs), up to 3none
RWKV-6Recurrent0about 33 MiB for the 7B model (FP32 state)
Ten isometric bars sorted from tall to flat, with coins in front of the hybrids for their fixed state.Point at a bar to name its model.
Every model type, sorted by cache per token, with fixed state shown separately.

Every number here is computed from public configurations, not measured. The BF16 bytes per token are model math, the same on every engine; DeepSeek-V3.2's FP8 row follows vLLM's sparse-MLA layout, and Nemotron-H's state is shown in BF16, where vLLM's default keeps it in FP32. Real engines add block padding, a memory reserve, CUDA graph buffers and, for FP8, scale tensors. Treat the table as the floor a measurement starts from. The data is available as CSV.

Why a hybrid model's prefix match is fragile

A prefix cache lets a new request reuse the cache of an earlier request that started with the same tokens, such as a long system prompt. The shared part is not computed again.

Two rows of blocks; the front row's first blocks are empty slots tied by dashed lines to the same blocks behind.Point at the two requests to find their shared blocks.
Two requests that start the same share their first blocks, and that part is not computed again.

A hybrid model cannot reuse its attention cache alone. It also needs the fixed state as it was at the end of the shared part. vLLM saves that state only at a few block boundaries, and its blocks for hybrid models are large: 528 tokens for Qwen3.5-4B in BF16. In a plain single-node vLLM, if the saved state is missing, the attention match is thrown away too; a KV connector that accepts partial hybrid hits can keep it.

Two rows of large blocks; where the shared part ends, a dark cube marks a saved state that is missing.Point at the blocks, then at the saved state.
On a hybrid model the blocks match, but the saved state is missing, so the attention match is lost too (single-node vLLM).

vLLM issue #45238 shows how sharp the geometry is. With Qwen3.5-4B, 528-token blocks and prompts of about 2,200 tokens, a June 2026 build kept saved states at tokens 1,584 and 2,112. With a 1,600-token shared prefix, the state at 1,584 fell inside the shared part, and 52 of 64 requests hit. With a 1,500-token shared prefix, both fell in the unique part, and none hit. vLLM has since added a saved state at the detected end of the shared prefix (#37898), and the issue's author re-measured on a later build: the shorter shared prefix still hit only 25.5 percent of prefix tokens, against 54.0 percent for the longer one. The total loss is gone; half the reuse is still lost, and nothing in vLLM's metrics shows why. --prefix-match-unit can add saved states inside a short prefix, but it is off by default.

Two rows of large blocks; a cube marks the one saved state, first inside the shared part, then just past it.Point at each prompt to see where its saved state falls.
Shorten the shared part by 100 tokens and the saved states land past it: on the June 2026 build, hits fell from 52 of 64 to 0 (vLLM #45238).

Prefix caching on hybrids also has a cost when nothing hits. vLLM issue #60008 measured 13.4 to 16.2 percent lower output throughput (medians of six rounds; the issue's title says 11 to 16) on one NVFP4 Mamba hybrid on four GB200 GPUs, with prefix caching on and no hits at all, on a vLLM build with changes not yet merged. Later comments report a prototype that closes most of the gap. Both issues were open when this essay was published.

Two bars: full throughput with caching off, and a shorter bar whose missing share is drawn in orange.Point at each bar to compare throughput.
With no hits at all, prefix caching cost 13.4 to 16.2 percent throughput on one hybrid setup (vLLM #60008).

Images add one more rule. Two requests with the same text and different images must not share a cache, so the cache key must include the image.

Two trays with the same row of text blocks and different dot images, each ending in a blue key tag with its own code.Point at the two images.
Same words, different images: the cache key must include the image.

What crosses the wire between prefill and decode

In disaggregated serving, one group of GPUs runs prefill and another runs decode, so the cache moves between them once per request. Its size is the per-token number times the prompt length, plus any fixed state.

Two servers joined by one cable, with a blue block on it whose length is the cache to send.Point at a model to send its cache.
What crosses the wire for an 8K-token prompt, model by model.

One request's cache and state at 8K tokens, BF16, and its transfer time at ideal line rate: 50 GB/s per 400 Gb/s NIC (400 GB/s for eight), 450 GB/s per direction for NVLink 4.

Model

At 8K tokens

One 400G NIC

Eight 400G NICs

NVLink 4

Llama-2-7B4.0 GiB86 ms10.7 ms9.5 ms
Qwen3-8B1.125 GiB24 ms3.0 ms2.7 ms
Qwen3-30B-A3B768 MiB16 ms2.0 ms1.8 ms
DeepSeek-V3549 MiB11.5 ms1.4 ms1.3 ms
gpt-oss-120b292.5 MiB6.1 ms0.8 ms0.7 ms
Qwen3-Next-80B-A3B229.7 MiB4.8 ms0.6 ms0.5 ms
Nemotron-H-8B177.4 MiB3.7 ms0.5 ms0.4 ms

Real transfers are slower than line rate.

The layout on each side matters as much as the size. When prefill and decode use different numbers of GPUs, KV heads must be regrouped in flight. For MLA, every decode GPU needs the whole row, so the same bytes are sent to each one.

Four GPUs behind send slabs on arcs to eight GPUs in front, two heads from each GPU onto one each.Point at each side to see how the cache is laid out.
Different GPU counts on each side regroup the heads in flight; MLA sends its whole row to every GPU.

When reloading beats recomputing

A cache that does not fit in GPU memory can move to host memory, a local SSD or a remote store, and come back when a request needs it.

Four blocks stepping down and away, each wider than the last, with a small blue block of cache on the second.Point at a tier to see what it holds.
Each tier holds more and returns it more slowly.

Moving it back is worth it only when it is faster than computing it again. Reload time is the cache size divided by the tier's bandwidth, plus a fixed cost to start the transfer. Recompute time is the prefill time for those tokens, which grows faster than the prompt, because attention work grows with the square of its length. So short prompts favor recomputing and long prompts favor reloading. A slower tier, or a larger cache per token, moves the crossover to longer prompts; an MLA or hybrid model, with a small cache per token, moves it to shorter ones. Where it falls depends on numbers that only your hardware gives.

Three pairs of bars, blue for reload and orange for recompute; the orange bar grows much faster with the prompt.Point at a prompt length to compare the two times.
Reload wins when moving the bytes is faster than computing them again.

For the same reason, sparse attention does not make offload cheaper. It reads fewer tokens, but an offloaded cache must carry all of them.

Questions to ask any new model's config

Before you size a deployment for a new model, answer these from its config.json and its paper:

  1. What attention type does each layer use: full, sliding window, linear, Mamba or none?
  2. How many KV heads does each attention layer have, and is the head dimension stated separately?
  3. Does it store a latent row (MLA) instead of separate keys and values?
  4. Which layers have a window, and how long is it?
  5. Is there a fixed state per request, and how large is it in the dtype the engine uses?
  6. Do any layers share a cache?
  7. At your tensor-parallel size, is the cache split, or copied to every GPU?

The KV cache calculator runs these numbers for your model, context, GPUs and link.

Ten isometric bars, a hypothetical 128K request per model, from a tall Llama-2-7B down to a flat RWKV-6, Qwen3-8B in blue.Point at a model to see what a 128K-token request would hold.
Back to the start: a hypothetical 128K-token request, every model type side by side, even those whose context is shorter.

Questions engineers ask

Does an FP8 KV cache halve memory?

It halves the bytes per token, from 2 to 1 per element. Scale tensors add a small amount, and accuracy needs its own check for each model.

Why does my engine report fewer tokens than the formula?

The formula ignores block padding, the memory reserve, activations and CUDA graph buffers. Use the formula as the upper bound and the engine's startup log as the measurement.

Is the KV cache the same as the prefix cache?

No. The KV cache holds one request's keys and values. A prefix cache keeps those blocks after the request ends, so a later request with the same start can reuse them.

References