Tool
KV cache calculator
How much memory the KV cache takes for one token, one request and one GPU, for grouped-query, latent, sliding-window and hybrid models. It is the companion to the essay KV cache size per token, which explains every number.
Updated
How it computes
For each layer that keeps every token, the cache grows by 2 × KV heads × head dimension × bytes per element. Latent attention (MLA) stores one row of latent plus position values per layer instead. Sliding-window layers stop growing at their window. Hybrid models add a fixed state per request that does not grow.
Tokens are rounded up to whole cache blocks. Under tensor parallelism, KV heads are split across GPUs until there are more GPUs than heads; a latent row is never split by tensor parallelism, so every GPU holds all of it (decode context parallelism, which this page does not model, shards it by token instead). Weights are divided evenly across the GPUs.
Requests that fit is the memory left after the weights, divided by one request's cache on one GPU. It is an upper bound: activations, CUDA graph buffers and FP8 scale tensors also take memory, so check it against the engine's own startup log.
The memory share follows the engine you choose. vLLM's gpu_memory_utilization is a share of the GPU's total memory, and weights and cache both come out of it; it defaults to 0.92 since v0.20 (0.9 before), and this page starts at the essay's 0.9. SGLang's mem_fraction_static is also a share of total memory, but SGLang computes it at startup from the GPU and its settings, so enter the value its log reports. TensorRT-LLM's free_gpu_memory_fraction (the --kv_cache_free_gpu_memory_fraction flag) is a share of the memory left after the weights, 0.9 by default.
Bytes per token in BF16 are model math and the same on every engine. Latent attention in FP8 follows vLLM: sparse MLA (with an indexer) uses the 656-byte fp8_ds_mla row, plain MLA a 576-byte FP8 row; other engines may store a different row. Fixed state is shown in BF16, but vLLM keeps the SSM state of some hybrids in FP32 by default (Nemotron-H and the Kimi KDA models; Qwen3.5 by its config), which roughly doubles it. Extra multi-token-prediction layers, which some models carry for speculative decoding (GLM-5.3 has one), are not counted. Link rates are the vendors' published line rates per direction: NVLink and PCIe are per GPU, so the GPUs of a group receive in parallel, and the NICs of a node add up to a total the group shares.
Every model at a glance
| Model | Type | Bytes per token | 8K tokens | 32K tokens |
|---|---|---|---|---|
| Qwen3-8B | GQA | 147,456 | 1.125 GiB | 4.5 GiB |
| Llama-2-7B | MHA | 524,288 | 4 GiB | 16 GiB |
| Falcon-7B | MQA | 8,192 | 64 MiB | 256 MiB |
| Qwen3-30B-A3B | GQA, mixture of experts | 98,304 | 768 MiB | 3 GiB |
| DeepSeek-V3 | MLA | 70,272 | 549 MiB | 2.145 GiB |
| DeepSeek-V3.2 | MLA, sparse attention | 48,068 | 376 MiB | 1.467 GiB |
| gpt-oss-120b | GQA, sliding window | 36,864 | 293 MiB | 1.129 GiB |
| Qwen3-Next-80B-A3B | Linear attention hybrid | 24,576 | 230 MiB | 806 MiB |
| Nemotron-H-8B | Mamba hybrid | 16,384 | 177 MiB | 561 MiB |
The same data, with layer counts and config links, is available as CSV.
Any model on Hugging Face
The model field searches Hugging Face as you type and reads the chosen model's config.json in your browser, at its current commit. The classifier that builds the site's snapshot names the class and computes the cache. A model whose cache it cannot compute exactly, such as one with compressed or chunked attention, is marked unsupported with the reason; it is never estimated.
Before Hugging Face answers, the field offers a snapshot of 76 models, read on , chosen from Hugging Face's download counts and recent releases. It is a starting list, not a ranking: any public model is a search away. By class: 34 KV-head attention (MHA, GQA or MQA), 4 sliding window, 3 MLA, 2 MLA sparse, 4 Mamba hybrid, 11 linear hybrid, 18 unsupported. The snapshot, with every number and config link, is available as CSV.