Debarshi Das

Tool

KV cache calculator

How much memory the KV cache takes for one token, one request and one GPU, for grouped-query, latent, sliding-window and hybrid models. It is the companion to the essay KV cache size per token, which explains every number.

Updated

Calculator

Model and request
The model whose config.json sets the cache shape, read from Hugging Face. Type a name, an organization, a class such as mla or mamba, or paste a repo id or link.

The serving engine whose memory rule sets the cache budget. vLLM and SGLang count a share of total GPU memory; TensorRT-LLM counts a share of the memory left after the weights.
The bytes per element the engine stores keys and values in, set by its KV cache dtype option. FP8 latent rows follow vLLM: 656 bytes for sparse MLA, 576 for plain MLA. Latent caches are not offered in FP4.
The tokens one request keeps: its prompt and its generated tokens. They are rounded up to whole cache blocks.
The requests decoding at the same time. The results say whether their cache fits beside the weights.
GPUs
The memory per GPU as the driver reports it, in GiB. An H100 80GB reports about 79.6 GiB.
The share of each GPU's memory the serving engine may use, for weights and cache together: vLLM's gpu_memory_utilization. vLLM's default is 0.92 since v0.20 (0.9 before); this page starts at the essay's 0.9.
The GPUs one copy of the model is split across. KV heads split across them until there are more GPUs than heads; a latent row stays whole on every GPU.
The bytes per parameter the weights are stored in. The parameter count is the safetensors total that Hugging Face publishes for the model.
The tokens per cache block, set by the engine's block size option. Each request's tokens are rounded up to whole blocks.
Handoff link
The link the cache crosses from prefill to decode, at its vendor's published rate per direction. NVLink and PCIe are per GPU, so the GPUs of a group receive in parallel.
NICs per node that carry the handoff. Their rates add up to a node total that the tensor-parallel group shares.
A link rate in GB/s per direction, for a fabric that is not listed. The group shares it, as it shares a node's NICs.

Architecture of Qwen3-8B

Class
GQA
Layers with attention
36 of 36
KV heads
8
Head dimension
128
Sliding window
none
Fixed state per request
none
Parameters
8,190,735,360

From the site's snapshot of : config.json at b968826.

The config fields this rests on
  • model_type=qwen3
  • num_hidden_layers=36
  • num_attention_heads=32
  • head_dim=128
  • num_key_value_heads=8

Qwen3-8B, 8,192 tokens, BF16 cache

Cache per token
147,456 bytes
One request, cache and state
1.125 GiB
One request, per GPU
1.125 GiB
Weights per GPU
15.3 GiB
Left for the cache per GPU
56.4 GiB
Requests that fit, at most
50
64 requests need, per GPU
72 GiB (does not fit)
Handoff, prefill to decode
24.2 ms (at 50 GB/s, shared by the group)

An upper bound: activations, CUDA graph buffers and FP8 scale tensors also use memory. The budget follows vLLM: gpu_memory_utilization is a share of total GPU memory, for weights and cache together. Handoff times assume ideal line rate.

How it computes

For each layer that keeps every token, the cache grows by 2 × KV heads × head dimension × bytes per element. Latent attention (MLA) stores one row of latent plus position values per layer instead. Sliding-window layers stop growing at their window. Hybrid models add a fixed state per request that does not grow.

Tokens are rounded up to whole cache blocks. Under tensor parallelism, KV heads are split across GPUs until there are more GPUs than heads; a latent row is never split by tensor parallelism, so every GPU holds all of it (decode context parallelism, which this page does not model, shards it by token instead). Weights are divided evenly across the GPUs.

Requests that fit is the memory left after the weights, divided by one request's cache on one GPU. It is an upper bound: activations, CUDA graph buffers and FP8 scale tensors also take memory, so check it against the engine's own startup log.

The memory share follows the engine you choose. vLLM's gpu_memory_utilization is a share of the GPU's total memory, and weights and cache both come out of it; it defaults to 0.92 since v0.20 (0.9 before), and this page starts at the essay's 0.9. SGLang's mem_fraction_static is also a share of total memory, but SGLang computes it at startup from the GPU and its settings, so enter the value its log reports. TensorRT-LLM's free_gpu_memory_fraction (the --kv_cache_free_gpu_memory_fraction flag) is a share of the memory left after the weights, 0.9 by default.

Bytes per token in BF16 are model math and the same on every engine. Latent attention in FP8 follows vLLM: sparse MLA (with an indexer) uses the 656-byte fp8_ds_mla row, plain MLA a 576-byte FP8 row; other engines may store a different row. Fixed state is shown in BF16, but vLLM keeps the SSM state of some hybrids in FP32 by default (Nemotron-H and the Kimi KDA models; Qwen3.5 by its config), which roughly doubles it. Extra multi-token-prediction layers, which some models carry for speculative decoding (GLM-5.3 has one), are not counted. Link rates are the vendors' published line rates per direction: NVLink and PCIe are per GPU, so the GPUs of a group receive in parallel, and the NICs of a node add up to a total the group shares.

Every model at a glance

Bytes per token and one request's cache and state, BF16 (DeepSeek-V3.2 in FP8), 16-token blocks. Computed from each model's config.json, not measured.
ModelTypeBytes per token8K tokens32K tokens
Qwen3-8BGQA147,4561.125 GiB4.5 GiB
Llama-2-7BMHA524,2884 GiB16 GiB
Falcon-7BMQA8,19264 MiB256 MiB
Qwen3-30B-A3BGQA, mixture of experts98,304768 MiB3 GiB
DeepSeek-V3MLA70,272549 MiB2.145 GiB
DeepSeek-V3.2MLA, sparse attention48,068376 MiB1.467 GiB
gpt-oss-120bGQA, sliding window36,864293 MiB1.129 GiB
Qwen3-Next-80B-A3BLinear attention hybrid24,576230 MiB806 MiB
Nemotron-H-8BMamba hybrid16,384177 MiB561 MiB

The same data, with layer counts and config links, is available as CSV.

Any model on Hugging Face

The model field searches Hugging Face as you type and reads the chosen model's config.json in your browser, at its current commit. The classifier that builds the site's snapshot names the class and computes the cache. A model whose cache it cannot compute exactly, such as one with compressed or chunked attention, is marked unsupported with the reason; it is never estimated.

Before Hugging Face answers, the field offers a snapshot of 76 models, read on , chosen from Hugging Face's download counts and recent releases. It is a starting list, not a ranking: any public model is a search away. By class: 34 KV-head attention (MHA, GQA or MQA), 4 sliding window, 3 MLA, 2 MLA sparse, 4 Mamba hybrid, 11 linear hybrid, 18 unsupported. The snapshot, with every number and config link, is available as CSV.