Prefill-decode disaggregation separates prompt processing from token generation so each LLM phase can use different workers, schedules, and capacity plans.
A long RAG prompt demands a large burst of compute before its first token, while decoding performs many smaller, memory-bandwidth-sensitive iterations afterward. Running both phases on one GPU is simple but lets long prefills interrupt active conversations. Disaggregation moves the request and its KV state between pools, trading extra coordination and transfer cost for independent control of first-token and per-token latency.
Prefill and Decode Have Different Resource Shapes
Prefill processes all prompt tokens in parallel and builds the KV cache, producing a compute-intensive burst whose duration grows with prompt length. Decode repeatedly reads model weights and accumulated KV state to generate one or a few new tokens.
DistServe identifies prefill-decode interference when both phases share GPUs and ties prefill to time-to-first-token while decode drives time per output token. Separating them lets the scheduler protect each objective independently. This distinction remains visible during later household testing.
Disaggregation is not ordinary model parallelism. The same model may exist in both pools, while requests move between functional phases rather than layers of one forward pass. The intermediate result must remain inspectable before automation follows.
KV Transfer Connects the Two Worker Pools
After prefill, the system must make the request's KV cache available to a decode worker. It may transfer tensors over PCIe or a network fabric, use shared memory, or place workers to minimize the movement cost.
Splitwise studies phase-specific serving with phase-specific machines and scheduling, showing why hardware allocation can match the different computational characteristics of prompt and token work. Queueing and state movement become part of the serving path. That boundary should be measured separately under realistic operating conditions.
The decode pool cannot start until it has consistent KV state and request metadata. Large contexts increase transfer bytes, so a nominally faster phase split may lose to colocated execution on a small home network.
Independent Scaling Changes Capacity Planning
Separate pools can add prefill capacity for long-document bursts without proportionally expanding decode capacity, or protect voice decoding while background summarization consumes prompt workers. Admission control can target two queues and two latency budgets. The practical consequence appears when several sources compete for limited context.
Mooncake describes a KV cache coordination that treats KV cache movement and storage as first-class serving concerns. The architecture illustrates that disaggregation shifts the bottleneck from pure GPU scheduling toward state transfer and cache coordination.
The failure boundary is insufficient scale or bandwidth. One or two home GPUs may have no spare device for specialization, and duplicated model weights plus KV transfer can consume more memory and latency than the interference being removed.
Compare Colocated and Disaggregated Phase Budgets
Measure prompt tokens per second, time to first token, time per output token, KV bytes transferred, transfer time, queue wait, model-memory duplication, energy, and failure recovery for short, long, and mixed prompts. This dependency should remain explicit in the final interface.
Use chunked prefill as the colocated alternative. Test chunked prefill before adding a second pool, then compare identical arrival traces under both architectures. The result must therefore be checked against the original evidence.
Adopt disaggregation only when phase interference is measured and the transfer path preserves both latency objectives. On a small server, colocated scheduling with bounded prefill chunks may deliver the same user outcome with less state movement.
Tech & AI HUB
More to Read

What Is Embedding Drift, and When Does a Private Search Index Need Rebuilding?
Decode model, preprocessing, corpus, and query drift; distinguish monitoring from incompatibility; and decide when a private index needs rebuilding.

What Is Tokenizer Compatibility, and Why Can It Break Model Switching?
Decode vocabulary identity, special-token semantics, chat templates, cached tokens, adapters, and compatibility checks for local model switching.

What Is Model Residency, and When Should a Local AI Service Keep Weights Loaded?
Decode weight residency, cache levels, cold starts, eviction, multiplexing, memory pressure, and when a home AI service should stay warm.

