Several local models can share one accelerator when the serving layer coordinates weight residency, dynamic memory, execution time, and request isolation across workloads.
A home GPU may alternate between a chat model, an embedding model, a vision encoder, and speech recognition. Loading every weight set permanently may exceed VRAM, while unloading on each request makes first-token latency erratic. A multi-model controller needs residency policy, cross-model memory accounting, scheduling, cache isolation, preemption, and fairness rather than relying on separate processes to compete blindly.
Residency Policy Decides Which Weights Stay Warm
The controller tracks model size, arrival rate, load time, latency objective, and recent use. Popular models remain resident, infrequent models occupy CPU or storage, and predicted demand can trigger prewarming before the next request reaches the accelerator.
multi-model prewarming prepares universal GPU workers for multiple models and coordinates prewarming with eviction-aware placement. Its results show why avoiding a cold load can dramatically improve time to first token under predictable demand. This distinction remains visible during later household testing.
Residency decisions should include quantization and adapter variants because two apparently similar endpoints may hold different base weights. A strict memory budget prevents proactive loading from evicting the KV cache of active requests. The intermediate result must remain inspectable before automation follows.
Cross-Model Memory Coordination Prevents Fragmented Capacity
Weights are mostly stable, while activations and KV caches expand with batch and sequence length. A shared allocator can map memory pages on demand, reclaim idle regions, and expose reservations so one model cannot consume space promised to another.
cross-model memory coordination introduces cross-model memory coordination with dynamic virtual-to-physical page mapping and runtime sharing policies. The design explains why ordinary process-level GPU sharing cannot respond well to rapidly changing model demand. That boundary should be measured separately under realistic operating conditions.
Memory sharing is not data sharing. KV-cache blocks, prefix caches, temporary buffers, and adapter state require tenant and model identifiers; otherwise a reused page or cache key can leak context or corrupt results across endpoints.
Scheduling and Adapter Multiplexing Control Execution Time
A scheduler chooses between spatial sharing, where models occupy memory together, and temporal sharing, where kernels take turns. Continuous batching improves throughput, while preemption and weighted queues protect an interactive request from a long background job.
adapter multiplexing serves thousands of low-rank adapters over shared base models by paging adapter weights and coordinating heterogeneous batches. It demonstrates how specialization can share more state than completely separate model replicas. The practical consequence appears when several sources compete for limited context.
The failure boundary is kernel and memory interference. Two models that fit simultaneously may still miss latency targets when they contend for compute, bandwidth, or copy engines. Sharing is useful only when per-model p95 latency and fairness remain inside policy, not when aggregate utilization merely looks high.
Build a Model-Sharing Interference Matrix
Measure each model alone, then run every important pair and the expected four-model mix with short, long, bursty, and background requests. Record cold-load time, resident memory, KV growth, kernel utilization, throughput, p50 and p95 latency, and evictions.
Use the routed-stack principle in routed model residency to assign priority and a residency class to each endpoint. Repeat with adapters, quantized variants, continuous batching, and preemption while checking that caches and request identities remain isolated.
Keep a sharing policy only when important interactive models meet their latency target during the worst expected mix. If one pair causes repeated thrashing, serialize that pair or reserve a time window instead of increasing concurrency for a better utilization graph.
Tech & AI HUB
More to Read

What Features Enable a Home AI Trust Boundary Around Sensitive Files?
See how classification, capability-scoped access, isolated parsing, retrieval filters, egress policy, approvals, and audits contain sensitive home files.

What Factors Determine Whether Merkle-Tree Backups Detect Silent Change Efficiently?
Learn how chunk size, fan-out, trusted roots, cached hashes, change locality, metadata scope, and scrubbing determine Merkle backup verification cost.

What Components Enable Verifiable Backups of AI Indexes and Model State?
See how coordinated snapshots, content manifests, checksums, version locks, restore drills, and query tests prove that AI state can actually recover.

