First-token latency usually spikes after model switching because the newly selected model must rebuild resident state before it can process the prompt.
A home AI server may answer quickly with one warm model, switch to a vision or coding model, and then pause before the first token. The delay can include weight reads, dequantization, device transfer, kernel compilation, CUDA graph capture, cache allocation, and prompt prefill. Which stage dominates depends on storage, available accelerator memory, runtime policy, and whether the prior model was fully evicted.
Cold Weight Loading Is the First Cause Family
If the next model is absent from RAM or VRAM, the server must read one or more checkpoint shards, validate or map them, construct tensors, and transfer usable weights to the execution device. A larger file or slower NAS path lengthens the pause before inference can begin.
Measurements of LLM cold-start latency show that startup can dominate TTFT when model state is cold. This symptom appears as heavy storage reads and model-loader time before any prompt-prefill kernels run. This distinction remains visible during later household testing.
If switching between two models that both remain resident produces the same spike, weight loading is not the complete explanation. The distinguishing observation is whether bytes read and resident-model state change with the slow request.
Runtime Initialization Creates a Second Cold Path
A loaded model can still be operationally cold. The runtime may initialize a device context, select kernels, compile shapes, capture graphs, allocate KV blocks, or build tokenizer and prompt-template caches on the first request after activation.
An engineering analysis of model streaming and warm-up separates storage streaming from initialization and warm-up. The stage signature is modest checkpoint I/O followed by compilation, allocation, or accelerator activity before prompt processing. The intermediate result must remain inspectable before automation follows.
Model shape, quantization backend, context limit, batch profiles, and driver state determine which artifacts can be reused. Switching back quickly may be fast if caches survived, while a memory-pressure eviction makes the same path cold again.
Queueing and Prefill Can Masquerade as Model-Load Delay
The switch request may wait behind model shutdown, memory reclamation, another user, or a long prompt. Once admitted, prefill processes every input token before decode, so larger histories increase TTFT without changing model-loading time. That boundary should be measured separately under realistic operating conditions.
The paged KV-cache admission serving design explains how active sequences consume paged KV blocks and how admission depends on available cache capacity. A switch that changes cache reservations can therefore alter queue time independently of checkpoint size.
The failure boundary is a warm, resident model with stable initialization and a latency spike that tracks prompt length or concurrency. In that case, model switching is only correlated; prefill or scheduling is the direct cause.
Split TTFT Into Load, Warm-Up, Queue, and Prefill
Replay fixed prompts while logging model eviction, checkpoint bytes, storage throughput, host-to-device transfer, device-context creation, kernel compilation, graph capture, KV allocation, queue wait, prefill duration, and first decode step on one monotonic clock. The practical consequence appears when several sources compete for limited context.
Compare the trace with model routing by memory, then test cold switch, immediate switch-back, resident dual-model routing, short prompt, and long prompt separately. Preserve sampling, client, and concurrency so only the intended state changes. This dependency should remain explicit in the final interface.
Assign the spike to the earliest expanding stage. Keep weights resident when loading dominates, persist compatible artifacts when initialization dominates, and change admission or context policy when queueing or prefill—not the switch itself—sets TTFT.
Tech & AI HUB
More to Read

What Causes WebSocket Reconnect Loops in a Remote Home AI Interface?
Diagnose WebSocket loops across handshake, proxy, authentication, heartbeat, network path, session recovery, and client backoff layers.

What Causes Backup Checksums to Mismatch After an Interrupted Transfer?
Trace checksum mismatches through source snapshots, chunk manifests, resume offsets, partial files, transforms, storage writes, and final verification.

What Causes Duplicate Household Entities in a Private Knowledge Graph?
Diagnose duplicate knowledge-graph nodes by separating extraction variants, identity keys, resolution thresholds, source lineage, and concurrent merges.

