Debarshi Das

Explainer

Disaggregated prefill and decode on NVIDIA GPUs

How disaggregated prefill and decode works on NVIDIA GPUs: what the KV-cache handoff costs over NVLink and RDMA, how Dynamo and NIXL move it, and when not to split.

LLM inference, Disaggregated serving, NVIDIA Dynamo, KV cache, GPU infrastructure

Disaggregated prefill and decode splits one inference request across two groups of GPUs. One group reads the prompt and builds the key-value (KV) cache. A second group receives that cache and generates the answer token by token. The design pays off when the latency it removes is worth more than the cost of moving the cache between the two groups. On NVIDIA systems, the network topology sets that cost, so the topology decides whether the split is worth doing.

This essay explains the trade in the order an infrastructure team meets it: why the two phases want different machines, what the handoff costs in bytes and milliseconds, how NVLink, InfiniBand, and RoCE change that cost, and which parts of the NVIDIA stack do the work. It also works through four consequences that get less attention than they deserve:

  • whether the handoff can hide behind prefill depends on one ratio that does not change with prompt length;
  • grouped-query and latent attention, not serving frameworks, made the handoff cheap enough to split;
  • the prefill-to-decode ratio is also a fan-in ratio, so incast decides tail latency on the network;
  • a failed decode worker turns into a burst of prefill work that the prefill pool must absorb.

Why prefill and decode want different machines

Every request to a large language model runs in two phases. Prefill processes the whole prompt at once. Because every prompt token is available at the start, the GPU runs large matrix multiplications that keep its tensor cores busy. Prefill determines the time to first token (TTFT).

Decode produces the output one token at a time. Each step reads the model weights and the growing KV cache to produce a single new token. That work moves far more bytes than it computes on, so memory bandwidth limits it. Decode determines the inter-token latency (ITL), which some tools call time per output token.

The usual summary is that prefill is compute-bound and decode is memory-bound. That summary is a good first approximation, but it is not a law. Large decode batches raise arithmetic intensity, short prompts make prefill cheap, and mixture-of-experts models and quantization move the balance again. Treat the summary as a reason to measure, not as a conclusion.

The practical problem appears when both phases share the same GPUs. A long prefill that arrives in the middle of a decode batch delays every token in that batch. Users then see output stall while someone else's long prompt is processed. Chunked prefill reduces the problem by splitting a long prompt into pieces and mixing them with decode steps. Disaggregation removes the interference by putting the phases on separate workers.

What disaggregation changes, and what it does not

Disaggregated serving gives each phase its own pool. Prefill workers can use a parallelism layout that minimizes TTFT. Decode workers can use a different layout and large batches that maximize tokens per second at a target ITL. The two pools scale independently.

One cabinet of GPU sleds, divided into two pools. The sleds above the split run prefill and the sleds below run decode. Where the split sits is a decision you now own, and the right place depends on your traffic.

The research behind the idea is recent. DistServe, published at OSDI 2024, placed the phases on separate GPUs and optimized for goodput: the request rate a system serves while it meets both its TTFT and its ITL targets. Its authors report serving up to 7.4 times more requests, or meeting a 12.6 times tighter latency target, than the systems they compared against. Splitwise, published at ISCA 2024, argued that the two phases can run on different types of hardware, and its authors report 1.4 times the throughput at 20 percent lower cost. Mooncake, the serving system behind Kimi, organized its whole design around the KV cache as the object that moves through the cluster. These are results their authors measured on their own systems, so read them as evidence that the idea works, not as a forecast for your fleet.

By late 2025, the DistServe authors wrote that almost every production-grade serving framework had adopted disaggregation in some form.

Disaggregation does not raise raw throughput by itself. The vLLM documentation says so plainly: "Disaggregated prefill DOES NOT improve throughput." Its stated benefits are separate tuning of TTFT and ITL, and tighter control of tail ITL. The honest goal is throughput at a latency target. A team that expects a free throughput gain will measure the wrong thing.

The handoff: how large the KV cache is

Every disaggregated request ends prefill with a KV cache that must reach a decode worker. The size of that cache follows directly from the model:

  • KV bytes per token = 2 × layers × KV heads × head dimension × bytes per element.

The factor of 2 counts the keys and the values. Llama 3.1 70B has 80 layers, 8 KV heads through grouped-query attention, and a head dimension of 128. In BF16 each element takes 2 bytes, so the cache costs 327,680 bytes per token, or 320 KiB. In FP8 it costs half that.

Prompt length turns that figure into the size of a single handoff:

  • A 2,048-token prompt produces about 0.67 GB of KV cache in BF16.
  • An 8,192-token prompt produces about 2.7 GB.
  • A 32,768-token prompt produces about 10.7 GB.
  • A 128,000-token prompt produces about 42 GB. The Mooncake project cites a matching figure of about 40 GB for 128k tokens on Llama 3 70B.

Now divide by bandwidth. The figures below are line-rate floors with no protocol overhead, so real transfers take longer:

  • One 400 Gb/s NIC moves about 50 GB/s. The 8,192-token cache takes about 54 ms.
  • Eight such NICs in parallel, one per GPU, bring that down to about 7 ms. A DGX B200, for example, has eight ConnectX-7 NICs at 400 Gb/s.
  • NVLink gives each Hopper GPU 900 GB/s of total bandwidth and each Blackwell GPU 1.8 TB/s. On Blackwell that is about 900 GB/s in each direction, so the same cache takes about 3 ms between one pair of GPUs.

Published measurements land in the same range. The Mooncake Transfer Engine reports up to 87 GB/s over four 200 Gb/s RoCE links and up to 190 GB/s over eight 400 Gb/s RoCE links. DistServe's paper gives an older example: a single 512-token request on OPT-66B carries about 1.13 GB of KV cache, so ten such requests per second need about 90 Gb/s of transfer bandwidth.

These numbers frame the whole decision. A transfer of 3 to 7 ms is small against a TTFT budget of a few hundred milliseconds. A transfer of 54 ms, or 215 ms for a 32,768-token prompt over one NIC, is not. The cost also grows with load, because many requests hand off at the same time and share the same links.

Three techniques shrink the visible cost. An FP8 KV cache halves the bytes. Layer-wise transfer sends the cache for each layer while prefill computes the next one, so most of the transfer overlaps with compute. Splitwise reports that the part left without overlap was about 5 ms on H100. Larger KV blocks also help: Dynamo recommends a block size of 128 tokens, because small blocks make the handoff slow.

A ratio that does not depend on prompt length

Layer-wise transfer raises a sharper question: can the handoff hide completely behind prefill? The answer turns out to be a property of the model and the hardware, not of the prompt.

Both costs grow with the number of prompt tokens. Prefill does roughly 2 floating-point operations per parameter per token, so its time per token is about 2 × parameters ÷ effective FLOPS. The handoff moves a fixed number of KV bytes per token, so its time per token is KV bytes per token ÷ bandwidth. Divide one by the other and the prompt length cancels out. Attention adds work that grows faster than linearly for long prompts, which only makes prefill slower and the ratio more favorable.

The ratio for a few real configurations, using public specifications. Prefill runs on eight H100 GPUs at 989 dense BF16 TFLOPS each, and I assume 50 percent model FLOPS utilization:

  • Llama 3 70B, BF16 cache, eight 400 Gb/s NICs: the transfer takes about 2 percent of prefill time.
  • The same model over a single NIC: about 18 percent.
  • DeepSeek-V3, BF16 cache, eight NICs: about 1 percent.

Below 100 percent, layer-wise transfer can in principle finish inside prefill, and the handoff adds almost nothing to TTFT. The planning question changes as a result. A team does not need a separate handoff budget for every prompt length. It needs one number per model and hardware pair, and a check that the transfer path is pipelined. This analysis is mine, and it ignores queueing and protocol overhead, which the benchmark section below covers.

Attention design made disaggregation practical

The same arithmetic shows why disaggregation arrived when it did. KV bytes per token depend on how the model stores attention, and that design changed:

  • Multi-head attention keeps keys and values for every head. A 70B-class model with 64 KV heads would need about 2.5 MiB per token.
  • Grouped-query attention, which Llama 3 uses, shares 8 KV heads among the 64 query heads. That cuts the cache to 320 KiB per token, eight times smaller.
  • Multi-head latent attention, which DeepSeek-V3 uses, stores one compressed vector of 576 values per layer. Its 61 layers come to about 69 KiB per token in BF16, close to 37 times smaller than the multi-head case.

Run the ratio from the previous section on the multi-head variant and the result changes completely. Over a single NIC, the transfer would take about 147 percent of prefill time, so it could never hide behind compute. The handoff would sit squarely on the critical path.

In other words, the papers and frameworks of 2024 and 2025 did not make disaggregation cheap on their own. Model architects had already shrunk the object that needs to move. Anyone who evaluates disaggregation for a new model should start with one figure: KV bytes per token.

On NVIDIA systems there are two very different places to put the cut between prefill and decode.

Inside an NVLink domain, GPUs read and write each other's memory at NVLink speed. In an HGX or DGX server, the domain is the eight GPUs in the box. In a GB200 NVL72 rack, NVIDIA connects 72 Blackwell GPUs in a single NVLink domain with 130 TB/s of aggregate bandwidth. A prefill group and a decode group that both sit inside one NVL72 domain can hand off the cache at NVLink speed, even though they span many trays. This is the main reason rack-scale NVLink matters for disaggregated serving.

Outside the NVLink domain, the handoff crosses the scale-out network. On NVIDIA systems that usually means InfiniBand or RoCE through ConnectX NICs. GPUDirect RDMA lets the NIC read the cache straight from GPU memory, with no copy through host memory. That path only works when the drivers and memory registration are right, and when the GPU and the NIC share the same upstream PCIe root complex, which NVIDIA's documentation lists as a requirement. NVIDIA's Dynamo setup guide warns that falling back to TCP over Ethernet makes this transfer 200 to 500 times slower.

DistServe made the same point with its placement algorithms. On clusters with a fast cross-node network, it places prefill and decode instances freely. On clusters without one, it keeps each prefill instance and its decode instance on the same node, so the handoff stays on NVLink. The topology decides the architecture as well as the speed.

Disaggregation also creates a traffic pattern that scale-out networks handle poorly: incast. With a ratio of several prefill workers to one decode worker, several senders can finish their prompts at the same moment and push their caches to the same decode node. Four 8,192-token Llama 3 70B caches arriving together carry about 10.7 GB. Even across eight 400 Gb/s NICs, that burst occupies the links for about 27 ms, and later arrivals queue behind it.

The prefill-to-decode ratio is therefore also a fan-in ratio for the network. On RoCE, incast is where congestion control, ECN marking, and PFC pause behavior decide the tail latency. On InfiniBand, credit-based flow control behaves differently but still queues. Size the fabric for the burst at the decode NIC, not for the average bandwidth across the cluster. Spreading handoffs across decode workers in the router helps as much as adding bandwidth. This is a second reason an NVL72 domain matters: within the rack, the NVLink switch fabric absorbs that fan-in without touching the scale-out network.

Incast at one decode port. Each prefill sled that finishes at the same moment adds a KV cache to the queue, and the last one waits for all the others. The time in the corner uses the 8,192-token Llama 3 70B cache and eight 400 Gb/s NICs from the paragraph above.

Some practical consequences follow:

  • Measure the bandwidth your application achieves, not the line rate on the switch.
  • Check that each GPU uses the NIC that sits closest to it on the PCIe tree. A mismatched mapping forces traffic through a slower path.
  • Treat InfiniBand and RoCE as configurations to measure. Achieved bandwidth, congestion behavior, and tail latency matter more than the name of the fabric.

Who does what in the NVIDIA stack

Many descriptions blur these layers together. They are easier to reason about as a stack, from the top down:

  • Orchestration and routing: NVIDIA Dynamo. Dynamo is an open-source distributed serving framework that sits above the inference engines. Its KV-aware router tracks which worker holds which KV blocks and sends each request where it will recompute the least. Its Planner scales the prefill pool on queued tokens and the decode pool on KV-cache use, against TTFT and ITL targets. Its KV Block Manager offloads blocks from GPU memory to host memory, local disk, and object storage. Dynamo does not run the model itself.
  • Inference engines: TensorRT-LLM, vLLM, and SGLang. The engine executes the model on each worker. Dynamo supports all three as backends.
  • KV connectors. The engine hands the finished cache to a connector. vLLM documents connectors that include NixlConnector, LMCacheConnectorV1, and MooncakeConnector. With NixlConnector, the decode worker pulls the KV blocks from the prefill worker over RDMA after a short handshake. SGLang offers Mooncake and NIXL as transfer engines, and TensorRT-LLM uses NIXL over UCX by default.
  • Transfer library: NIXL. The NVIDIA Inference Xfer Library is "targeted for accelerating point to point communications in AI inference frameworks." It gives one interface over GPU memory, CPU memory, and storage. It reaches them through plugins such as UCX, GPUDirect Storage, and libfabric, and it picks a backend from the source and destination memory types.
  • Transport and hardware. Underneath, the bytes move over NVLink, or over InfiniBand or RoCE with GPUDirect RDMA. On NVL72 racks, SGLang's Mooncake engine can also transfer over multi-node NVLink.
The six layers, from the link at the bottom to Dynamo at the top. Dynamo routes and plans, the engine runs the model, the connector hands off the cache, NIXL chooses how it moves, and the transport and the link carry the bytes.

One constraint runs through every layer. Dynamo's documentation requires the prefill and decode workers to use the same model, data type, block size, and KV layout, because the decode worker reads the blocks exactly as prefill wrote them.

Two definitions prevent most of the confusion. NIXL is the transfer abstraction that picks a transport for each move. Dynamo is the layer that coordinates the engines that run the model.

Routing, prefix reuse, and multi-turn conversations

The first request in a conversation is simple: prefill builds the cache, decode receives it, and decode produces the answer. The second turn is harder. The most useful cache now lives on the decode worker, because it holds the whole conversation so far. A naive design sends the next turn back to a prefill worker that recomputes the shared prefix from nothing.

A decode worker that holds one KV cache for each conversation. A KV-aware router sends the next turn to the cartridge that already holds its history, so no prefill worker has to recompute it. A taller cartridge holds a longer history and saves more work.

This is where KV-aware routing earns its place. A router that knows which workers hold which prefixes can send a request where its cache already lives, or move only the missing part. Prefix reuse also matters for system prompts that thousands of requests share. Measure the prefix-cache hit rate before and after any change to routing, because a split that destroys cache locality can lose more than it gains.

Sizing the prefill and decode pools

There is no universal ratio between prefill and decode workers. The right ratio depends on the traffic:

  • Input sequence length. Long prompts need more prefill capacity.
  • Output sequence length. Long answers need more decode capacity.
  • Arrival rate and burstiness. These set how much headroom each pool needs.
  • Latency targets. A strict TTFT target needs spare prefill capacity. A strict ITL target limits decode batch size.

Each pool can also use its own parallelism. Prefill often benefits from more tensor parallelism to cut TTFT. Decode often benefits from more replicas with large batches. The engine and connector must support the combination you choose, so read the compatibility notes for your model architecture, KV-cache data type, and block size before you design the fleet. Because traffic shifts during the day, a planner that rebalances workers between pools is more useful than a fixed ratio chosen once.

Failure: a lost decode worker becomes a prefill storm

Disaggregation also changes what a failure costs. A decode worker holds the KV cache for every sequence it serves. When it fails, those caches disappear. Each affected request must either fail or return to the prefill pool and rebuild its cache from the full prompt and the tokens generated so far.

That recovery traffic arrives all at once and lands on the pool sized for normal arrivals. Take a decode worker that serves 300 conversations with 8,192 tokens of context each. Rebuilding those caches means prefilling about 2.5 million tokens. At the rate estimated above for Llama 3 70B on eight H100 GPUs, that is close to 90 seconds of work for one prefill worker, released in a single moment. New requests miss their TTFT targets while the backlog clears. Offloading KV blocks to host memory or storage, as Dynamo's KV Block Manager can, gives the system a copy to restore instead of recomputing. Either way, size prefill headroom for steady traffic plus the largest decode worker you can lose. Planners that scale down without draining in-flight requests make the same storm happen on purpose, so check how yours retires a worker.

How to benchmark disaggregated serving

A useful benchmark compares disaggregated serving against a well-tuned aggregated baseline with chunked prefill enabled, on the same hardware and the same traffic. Report at least these measurements:

  • TTFT at p50, p95, and p99.
  • ITL at p50, p95, and p99.
  • Goodput: the request rate served within both latency targets.
  • Achieved KV-transfer bandwidth in GB/s, and transfer time per request.
  • Queueing delay in each pool.
  • GPU utilization in each pool.
  • Prefix-cache hit rate.

Use a realistic mix of prompt and output lengths. A benchmark with one fixed prompt length hides the case where disaggregation matters most, which is a long prompt that arrives while many short conversations are decoding.

What the published numbers say

The headline results for disaggregated serving come mostly from vendors, and each one depends on its conditions. Read them as upper bounds measured on favorable workloads:

  • At launch in March 2025, NVIDIA reported that Dynamo served up to 30 times more requests for DeepSeek-R1 on GB200 NVL72, with TensorRT-LLM, FP4 weights, and 32K input and 8K output tokens. For Llama 70B on Hopper, with vLLM and FP8 at 3K input and 50 output tokens, it reported more than twice the throughput.
  • A June 2025 NVIDIA blog reported up to 6 times the throughput for DeepSeek-R1 on GB200 NVL72. Those figures came from a GPU performance simulator, not from measured runs.
  • The TensorRT-LLM team reported gains of 1.4 to 1.8 times for DeepSeek-R1 on GB200 at 4,400 input and 1,200 output tokens, rising with multi-token prediction.
  • SemiAnalysis, an independent analyst, measured DeepSeek-R1 in FP4 with Dynamo and TensorRT-LLM. At 125 tokens per second per user, GB200 NVL72 produced 4.39 times the throughput per GPU of a B200 system, which it attributes largely to the larger NVLink domain.

The pattern is consistent. The largest gains appear with long prompts, large mixture-of-experts models, and a big NVLink domain. Those are exactly the conditions where the handoff is cheap and the interference is expensive.

When not to disaggregate

Treat disaggregation as an option to measure. NVIDIA's Dynamo documentation puts it directly: "For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster." Keep an aggregated deployment when:

  • Prompts are short, so prefill rarely disrupts decode. Dynamo's tuning guide notes that prefill runs inefficiently below roughly 1,000 input tokens.
  • Load is low, so there is little interference to remove.
  • Chunked prefill already keeps tail ITL inside your target.
  • The handoff must cross a slow or poorly configured network.
  • The team cannot yet operate two pools, a router, and a transfer path, and debug failures across all of them.

The rule that holds the essay together is simple. Disaggregated serving is worth it when moving the KV cache costs less than letting prefill and decode interfere with each other. On NVIDIA hardware, the topology sets the price of that move.

Questions engineers ask

What is the difference between prefill and decode?

Prefill processes the entire prompt in one parallel pass and builds the KV cache. Decode then generates the output one token at a time from that cache. Prefill sets the time to first token, and decode sets the time between tokens.

Is disaggregated inference the same as disaggregated serving?

In practice, yes. Both terms describe running prefill and decode on separate workers and moving the KV cache between them. Papers often call it P/D disaggregation.

Does disaggregation need InfiniBand?

No. It needs a fast path for the KV cache. NVLink is the fastest path inside a server or an NVL72 rack. Across racks, InfiniBand and RoCE both work when GPUDirect RDMA is configured correctly. What matters is the bandwidth you achieve and how stable it is under load.

Is chunked prefill an alternative?

Yes. Chunked prefill keeps both phases on the same GPUs and interleaves them. It is simpler to operate and often good enough. Disaggregation gives stronger isolation of tail latency at the cost of a transfer path and a second pool.

References