So I would like to share my analysis of PD disagg that I've run by Horace and many folks at Thinky
You may have heard that PD disagg helps with per-token "tail latencies" in ITL/TPOT, as prefill interrupts decode. This never was a convincing explanation to me though because people don't actually see single-token tail latencies during decode (what matters is how long the entire response takes - 100-10k tokens!).
Insight: steady-state response tells us
1. PD disagg actually improves mean per-token latency
2. PD disaggregation helps most for prefill-heavy workloads, it doesn't help for ultra low-latency
[Image 1] Here's an intuitive model. Imagine having 4 regular samplers, getting traffic so they spend 75% of their time on prefill and 25% on decode.
The batch size (running requests) of each sampler is determined by a steady-state feedback loop. As new requests come in, they increase batch size, which then decreases interactivity as requests finish slower. If a decode batch takes 10ms, average TPOT would be 40ms in this case (only 25% time spent on decode).
[Image 2] Now let’s say that we split it into 3P/1D disaggregated samplers. Prefill takes the same amount of time, and 1 decode worker handles 4x the number of requests. But it runs with the same batch size (throughput is equal), as it processes 4x the number of decode requests, each 4x faster. Each decode batch would take 10ms, but average TPOT has been lowered from 40ms → 10ms.
In other words, the steady-state model of PD disagg is: if your regular engine is spending a substantial amount of time on prefill, then at scale, you can get "increased interactivity” with multiple 1/(Decode Time %), while keeping the exact same batch size and throughput. With 1:1 PD, you get 1/0.5 = 2x interactivity. With 1:4 PD, you get 1/0.8 = 1.25x interactivity.
I've validated this in production on a bunch of different inference workloads.
Basically - decode-bound (low-latency) inference does not need disagg. But if you see your inference engine is spending a good % of its time running prefill passes, it's time to bust out the PD disagg.
---
Assumptions (if you want to check my work):
- This is independent of any prefill delayer settings; it actually assumes "perfect" prefill/decode scheduling on the regular engines.
- This assumes that the network fabric for sending P->D kv caches is not your bottleneck - e.g., 100 GB/s Infiniband can transfer KV caches in 100 ms.
- Obviously if you do PD disagg, you should also specialize engine configurations - like prefill may use EP4 while decode uses wider EP32 topologies. This gives you additional speedup! But this is secondary.
- I don't include mixed chunk in this analysis because mixed-chunk has not very good support, e.g., it's incompatible with speculative decoding (very important optimization) in most engines.
- P/D disagg also affects the amount of usable GPU HBM you get for storing persistent KV prefix caches -- in this analysis you lose total KV space if prefill takes <50% of your time. I think this is largely mitigated by tiered caches + Mooncake in practice.
I find that surprisingly few people have *really* thought about why PD disagg is a good idea. Even fewer people have *really* thought about how PD disagg compares to mixed chunk prefill (as Matthew describes here).