Local model-server memory usually creeps because caches and allocators retain reusable blocks, though unbounded references or native leaks can cause true growth.
After each home-server request, a dashboard may show RAM or VRAM rising without returning to the original baseline. The runtime can keep KV blocks, prefix entries, kernels, graphs, workspaces, and freed tensor blocks for reuse. Variable prompt lengths can fragment pools, while logging, sessions, image buffers, or extensions retain objects indefinitely, so these mechanisms require different evidence and limits.
Caching Allocators Reserve Freed Blocks for Reuse
GPU frameworks avoid expensive device allocations by keeping released blocks in a process-owned pool. Application tensors can be gone while the driver still reports the reserved pool as used by the model server. This distinction remains visible during later household testing.
An allocator analysis explains how cached GPU memory blocks rounds, splits, merges, and caches CUDA blocks. The signature is allocated tensor memory falling after a request while reserved memory stays high and later requests reuse it.
This plateau is not automatically a leak. It becomes harmful when the pool prevents another service from allocating or keeps expanding under repeated requests of identical shape after warm-up. The intermediate result must remain inspectable before automation follows.
Serving Caches and Request Shapes Expand the Intended Working Set
KV caches grow with active context, prefix caches retain reusable prompts, and compiled graphs or kernels cover observed batch shapes. New context lengths, modalities, and concurrency profiles can add entries between requests. That boundary should be measured separately under realistic operating conditions.
Research on LLM memory fragmentation identifies fragmentation between activation and KV-cache memory spaces in LLM serving. The observation explains why total capacity can rise even when no single live request is large. The practical consequence appears when several sources compete for limited context.
Record cache entry counts and shape classes. Growth that stops after the workload distribution stabilizes is bounded warm-up; growth proportional to total request count or unique session IDs suggests missing eviction. This dependency should remain explicit in the final interface.
Retained CPU Objects and Native Buffers Produce True Creep
Request histories, streaming queues, metrics labels, tokenizer outputs, uploaded images, pinned host buffers, and extension allocations can remain referenced after completion. GPU snapshots may look stable while process RSS continues rising. The result must therefore be checked against the original evidence.
A practical investigation of allocated versus reserved memory separates allocated, reserved, and process memory signals. That layered view prevents a CPU-side retention problem from being misdiagnosed as GPU allocator behavior. This distinction remains visible during later household testing.
The failure boundary is a one-time rise followed by a stable high-water mark. Call it a leak only when controlled identical requests produce continuing retained growth after cache limits, garbage collection, and expected pools are accounted for.
Build a Per-Request Memory Retention Curve
Replay hundreds of identical requests, then mixed lengths and modalities, while recording GPU allocated and reserved bytes, KV and prefix entries, graph cache, pinned memory, process RSS, object counts, request sessions, worker restarts, and allocator snapshots.
Use post-request memory reservation to distinguish deliberate post-request reservation. Repeat with each optional cache, extension, upload path, and metrics label disabled separately while holding the model and concurrency fixed. The intermediate result must remain inspectable before automation follows.
Accept bounded warm-up that plateaus within the declared memory budget. Add eviction when cache cardinality grows without benefit, normalize request shapes when fragmentation dominates, and isolate a true leak only after retained allocation stacks point to an owner.
Tech & AI HUB
More to Read

What Causes WebSocket Reconnect Loops in a Remote Home AI Interface?
Diagnose WebSocket loops across handshake, proxy, authentication, heartbeat, network path, session recovery, and client backoff layers.

What Causes Backup Checksums to Mismatch After an Interrupted Transfer?
Trace checksum mismatches through source snapshots, chunk manifests, resume offsets, partial files, transforms, storage writes, and final verification.

What Causes Duplicate Household Entities in a Private Knowledge Graph?
Diagnose duplicate knowledge-graph nodes by separating extraction variants, identity keys, resolution thresholds, source lineage, and concurrent merges.

