Why Does ZFS ARC Shrink When Local AI Uses Pinned Host Memory?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

ZFS ARC shrinks during pinned host-memory use because unreclaimable AI transfer buffers increase pressure on memory the kernel can reclaim, including filesystem cache.

GPU runtimes pin host pages so devices can transfer data without those pages moving or being swapped mid-operation. That improves DMA predictability, but pinned allocations are poor reclaim targets when available RAM falls. Linux invokes reclaim and shrinkers, and ZFS responds by evicting cached blocks or lowering its ARC target, leaving more memory for allocations that cannot yield.

Pinned Pages Change Which Memory the Kernel Can Reclaim

Ordinary anonymous pages may be swapped and clean file-cache pages may be discarded. Long-term pinned pages remain resident because a device or driver depends on their physical mapping, reducing the flexible pool available to satisfy new allocations.

Linux documentation for long-term page pinning distinguishes long-term page pinning from ordinary references and explains why DMA users must mark pages appropriately. The mechanism makes pinned memory qualitatively different from a process allocation that the kernel can readily move or reclaim.

AI frameworks use pinned staging buffers for faster host-to-device copies, dataloader queues, and offload. Multiple workers or oversized prefetch queues can keep far more host memory pinned than one visible batch suggests. This distinction remains visible during later household testing.

ARC Is a Large Reclaimable Consumer by Design

The Adaptive Replacement Cache retains recently and frequently used ZFS blocks to avoid storage reads. On Linux it participates in memory-pressure handling and can reduce its resident size when the system needs pages elsewhere. The intermediate result must remain inspectable before automation follows.

The OpenZFS ARC size and reclaim documentation describes ARC size controls and reclaim-related tunables. A configured maximum is a ceiling, not a promise that cached data will remain resident under pressure. That boundary should be measured separately under realistic operating conditions.

When pinned buffers grow, ARC eviction may be the correct response rather than a leak. The consequence appears later as lower cache-hit ratios, more disk reads, and slower file access after the AI job ends until the cache warms again.

Unified Memory and Container Metrics Can Hide the Competition

On an integrated GPU, model tensors and filesystem cache draw from the same physical RAM even when dashboards label usage differently. On a discrete GPU, host staging remains separate from VRAM but still competes with ARC on the server.

The OpenZFS implementation of the Linux ARC shrinker implementation registers ARC reclaim behavior with kernel memory management. Source-level evidence helps distinguish deliberate cache shrinking from an application directly commanding ZFS to discard blocks. The practical consequence appears when several sources compete for limited context.

The failure boundary is assuming pinned memory whenever ARC falls. A large file scan, explicit ARC limits, metadata pressure, cgroup reclaim, virtual-machine growth, or normal adaptive behavior can produce the same graph. Confirm pin counts and allocation timing.

-15% OFF
Single board computer zimaboard2

Correlate Pinned Bytes With ARC Reclaim and Cache Misses

Run a fixed AI workload while recording pinned or unevictable memory, MemAvailable, reclaim stalls, ARC size and target, ARC hit ratio, ZFS reads, pinned-buffer pool size, worker count, batch size, VRAM, and request latency. Include a no-AI storage baseline.

Use shared host memory to separate CPU, memory, and storage symptoms. Repeat with pageable transfers, smaller prefetch depth, fewer workers, and a bounded pinned pool, changing only one control per run. This dependency should remain explicit in the final interface.

Call pinning causal when ARC contraction tracks pinned growth and eases under a smaller pool. Cap the pool if storage misses harm other services, but retain enough pinning to avoid starving the accelerator; the correct balance depends on concurrent NAS demand.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.