What Is Prefill-Decode Disaggregation, and Why Does It Change LLM Serving?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

Prefill-decode disaggregation separates prompt processing from token generation so each LLM phase can use different workers, schedules, and capacity plans.

A long RAG prompt demands a large burst of compute before its first token, while decoding performs many smaller, memory-bandwidth-sensitive iterations afterward. Running both phases on one GPU is simple but lets long prefills interrupt active conversations. Disaggregation moves the request and its KV state between pools, trading extra coordination and transfer cost for independent control of first-token and per-token latency.

Prefill and Decode Have Different Resource Shapes

Prefill processes all prompt tokens in parallel and builds the KV cache, producing a compute-intensive burst whose duration grows with prompt length. Decode repeatedly reads model weights and accumulated KV state to generate one or a few new tokens.

DistServe identifies prefill-decode interference when both phases share GPUs and ties prefill to time-to-first-token while decode drives time per output token. Separating them lets the scheduler protect each objective independently. This distinction remains visible during later household testing.

Disaggregation is not ordinary model parallelism. The same model may exist in both pools, while requests move between functional phases rather than layers of one forward pass. The intermediate result must remain inspectable before automation follows.

KV Transfer Connects the Two Worker Pools

After prefill, the system must make the request's KV cache available to a decode worker. It may transfer tensors over PCIe or a network fabric, use shared memory, or place workers to minimize the movement cost.

Splitwise studies phase-specific serving with phase-specific machines and scheduling, showing why hardware allocation can match the different computational characteristics of prompt and token work. Queueing and state movement become part of the serving path. That boundary should be measured separately under realistic operating conditions.

The decode pool cannot start until it has consistent KV state and request metadata. Large contexts increase transfer bytes, so a nominally faster phase split may lose to colocated execution on a small home network.

Independent Scaling Changes Capacity Planning

Separate pools can add prefill capacity for long-document bursts without proportionally expanding decode capacity, or protect voice decoding while background summarization consumes prompt workers. Admission control can target two queues and two latency budgets. The practical consequence appears when several sources compete for limited context.

Mooncake describes a KV cache coordination that treats KV cache movement and storage as first-class serving concerns. The architecture illustrates that disaggregation shifts the bottleneck from pure GPU scheduling toward state transfer and cache coordination.

The failure boundary is insufficient scale or bandwidth. One or two home GPUs may have no spare device for specialization, and duplicated model weights plus KV transfer can consume more memory and latency than the interference being removed.

-15% OFF
Single board computer zimaboard2

Compare Colocated and Disaggregated Phase Budgets

Measure prompt tokens per second, time to first token, time per output token, KV bytes transferred, transfer time, queue wait, model-memory duplication, energy, and failure recovery for short, long, and mixed prompts. This dependency should remain explicit in the final interface.

Use chunked prefill as the colocated alternative. Test chunked prefill before adding a second pool, then compare identical arrival traces under both architectures. The result must therefore be checked against the original evidence.

Adopt disaggregation only when phase interference is measured and the transfer path preserves both latency objectives. On a small server, colocated scheduling with bounded prefill chunks may deliver the same user outcome with less state movement.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.