Local LLM responses become shorter under load when the serving layer trades generation budget for concurrency through caps, deadlines, preemption, or failed requests.
The model does not inherently decide to be concise because another user arrived. Under a fixed prompt and sampling state, concurrency should mainly change queue and token timing. Shorter answers indicate that the runtime, gateway, client, or memory manager changed an effective stopping condition, cancelled work, or returned a partial stream after pressure crossed a threshold.
Concurrency Expands KV Memory and Activates Serving Limits
Every active sequence holds KV-cache blocks that grow with retained context and generated tokens. When several requests share one accelerator, the runtime may lower maximum output, reject admission, preempt a sequence, or swap blocks to keep the batch within memory.
A serving design based on paged KV-cache allocation uses paged KV blocks to reduce fragmentation and enable higher concurrency. Its mechanism improves capacity, yet it also makes clear that each live sequence consumes a growing memory allocation until completion or eviction.
A gateway can impose a separate per-request or global token budget. If that budget is derived from available capacity, priority, or queue depth, identical prompts receive different maximum output even though the model weights and sampling parameters appear unchanged.
Deadlines and Preemption Can Return a Valid-Looking Partial Answer
Interactive systems often enforce wall-clock deadlines, idle-stream timeouts, or client cancellation. Slower inter-token delivery under load reaches those limits earlier in the semantic answer, and some APIs return the tokens already emitted instead of a conspicuous error.
The chunked prefill method splits prompt processing into smaller chunks to prevent long prefills from blocking decode. This work demonstrates how scheduling changes time to first token and inter-token latency under mixed request pressure. This distinction remains visible during later household testing.
Preemption may preserve a request for later resumption, restart it, or abort it depending on the engine. If the client disconnects during the pause, the server can log cancellation while the interface displays a grammatical but incomplete prefix as a finished response.
Sampling Alone Should Not Correlate Reliably With Load
Stochastic decoding naturally produces variable lengths when temperature and random seed differ. That variation can coincide with load in small samples, but concurrency has no direct semantic signal unless shared state, adaptive policy, or a software defect changes the decoding path.
Research on SLO-aware scheduling models routing and scheduling while protecting time-between-token objectives. The separation between throughput, TTFT, and decode deadlines shows why capacity policy must be measured independently from model output quality. The intermediate result must remain inspectable before automation follows.
The failure boundary is blaming the scheduler before checking stop metadata. End-of-sequence tokens, explicit length caps, client cancellations, server deadlines, OOM errors, and transport disconnects are different causes. Only repeated, controlled length shifts with matching stop reasons support a load mechanism.
Compare Length and Stop Reason at Fixed Concurrency Steps
Replay fixed prompts and seeds at one, two, four, and eight concurrent requests. Record requested max tokens, actual output tokens, finish reason, queue time, TTFT, inter-token latency, wall-clock deadline, client disconnect, preemption count, KV bytes, free VRAM, and server error.
Relate the memory behavior to concurrent workload limits, then repeat without gateway timeouts and with a fixed admission limit. Preserve prompts, templates, sampling, and client code so the only intended change is concurrency. That boundary should be measured separately under realistic operating conditions.
Treat shorter output as a serving defect when completion rate or semantic coverage drops before the advertised budget. If latency alone rises while finish reasons remain EOS, gather more seeded trials; if timeouts or caps dominate, expose and size those policies explicitly.
Tech & AI HUB
More to Read

Why Do Face Search Clusters Separate After Photos Are Rotated?
See how orientation metadata, pixel rotation, face alignment, crop geometry, resampling, and quality thresholds split face-search clusters after rotation.

Why Do Smart Home Charts Jump When Sensor Timestamps Are Rounded?
Learn how rounding, time buckets, aggregation, interpolation, clock zones, and duplicate timestamps create artificial jumps in smart home charts.

Why Do Agent Approval Prompts Reappear After the Browser Refreshes?
Learn how page state, session storage, server workflow records, approval scope, idempotency, and expiry make agent prompts reappear after refresh.

