What Causes First-Token Latency Spikes After a Local AI Service Switches Models?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

First-token latency usually spikes after model switching because the newly selected model must rebuild resident state before it can process the prompt.

A home AI server may answer quickly with one warm model, switch to a vision or coding model, and then pause before the first token. The delay can include weight reads, dequantization, device transfer, kernel compilation, CUDA graph capture, cache allocation, and prompt prefill. Which stage dominates depends on storage, available accelerator memory, runtime policy, and whether the prior model was fully evicted.

Cold Weight Loading Is the First Cause Family

If the next model is absent from RAM or VRAM, the server must read one or more checkpoint shards, validate or map them, construct tensors, and transfer usable weights to the execution device. A larger file or slower NAS path lengthens the pause before inference can begin.

Measurements of LLM cold-start latency show that startup can dominate TTFT when model state is cold. This symptom appears as heavy storage reads and model-loader time before any prompt-prefill kernels run. This distinction remains visible during later household testing.

If switching between two models that both remain resident produces the same spike, weight loading is not the complete explanation. The distinguishing observation is whether bytes read and resident-model state change with the slow request.

Runtime Initialization Creates a Second Cold Path

A loaded model can still be operationally cold. The runtime may initialize a device context, select kernels, compile shapes, capture graphs, allocate KV blocks, or build tokenizer and prompt-template caches on the first request after activation.

An engineering analysis of model streaming and warm-up separates storage streaming from initialization and warm-up. The stage signature is modest checkpoint I/O followed by compilation, allocation, or accelerator activity before prompt processing. The intermediate result must remain inspectable before automation follows.

Model shape, quantization backend, context limit, batch profiles, and driver state determine which artifacts can be reused. Switching back quickly may be fast if caches survived, while a memory-pressure eviction makes the same path cold again.

Queueing and Prefill Can Masquerade as Model-Load Delay

The switch request may wait behind model shutdown, memory reclamation, another user, or a long prompt. Once admitted, prefill processes every input token before decode, so larger histories increase TTFT without changing model-loading time. That boundary should be measured separately under realistic operating conditions.

The paged KV-cache admission serving design explains how active sequences consume paged KV blocks and how admission depends on available cache capacity. A switch that changes cache reservations can therefore alter queue time independently of checkpoint size.

The failure boundary is a warm, resident model with stable initialization and a latency spike that tracks prompt length or concurrency. In that case, model switching is only correlated; prefill or scheduling is the direct cause.

Split TTFT Into Load, Warm-Up, Queue, and Prefill

Replay fixed prompts while logging model eviction, checkpoint bytes, storage throughput, host-to-device transfer, device-context creation, kernel compilation, graph capture, KV allocation, queue wait, prefill duration, and first decode step on one monotonic clock. The practical consequence appears when several sources compete for limited context.

Compare the trace with model routing by memory, then test cold switch, immediate switch-back, resident dual-model routing, short prompt, and long prompt separately. Preserve sampling, client, and concurrency so only the intended state changes. This dependency should remain explicit in the final interface.

Assign the spike to the earliest expanding stage. Keep weights resident when loading dominates, persist compatible artifacts when initialization dominates, and change admission or context policy when queueing or prefillโ€”not the switch itselfโ€”sets TTFT.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.