How to Store Local LLM Models Without Filling the Workstation SSD

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

Keep a small hot set on workstation NVMe, place the larger model library on shared storage, and make every runtime use explicit paths.

This layout works when model weights are mostly read during startup, the network can deliver acceptable load times, and irreplaceable fine-tunes are protected separately from downloadable files. It fails when every cache silently falls back to the boot SSD or when a missing NAS makes the inference service redownload models into a new local directory.

Classify Model Files by Working Set and Rebuild Cost

Start with an inventory rather than moving one enormous cache directory. Base weights and quantized variants may be downloadable again, but adapters, fine-tunes, prompt templates, manifests, evaluation results, and locally converted files can be unique. Mark each item as hot, warm, cold, or irreplaceable, then record which application owns its path.

The hot set contains models used every day and should fit inside a fixed workstation budget. Warm models can live on the NAS and be copied locally before a project. Cold experiments can remain only in the shared library. Irreplaceable outputs need versioned backup even if their parent model can be downloaded again.

This classification prevents two common mistakes: backing up hundreds of gigabytes that are easy to recreate, and deleting a small adapter or manifest that cannot be recreated cheaply. It also provides the first capacity number: hot-set size plus free space for one incoming model, not the size of every model you may ever test.

Assign Local NVMe, Shared Storage, and Archive Roles

A real-world local AI cluster stored model files on a NAS and loaded them over 10GbE, showing that the pattern is viable when the network and storage path are designed for large reads. The useful lesson from that NAS-backed model-serving workflow is role separation: the shared library is the source, while compute and memory remain on the inference node.

Storage role Recommended contents Failure behavior Control
Workstation NVMe hot tier Current models, tokenizer files, active runtime cache Inference continues if NAS is unavailable Hard size quota and least-recently-used cleanup
NAS model library Approved weights, quantizations, shared revisions New loads pause; active in-memory model may continue Read-mostly share and checksums
Protected project store Fine-tunes, adapters, manifests, evaluation results Rebuild depends on backup Snapshots plus independent backup
Scratch space Partial downloads, conversions, temporary shards Safe to delete Separate path with automatic expiry

Do not point every runtime at the same writable network folder. A failed conversion, cleanup job, or version change could alter files used by another tool. Keep the canonical library read-mostly, stage changes in scratch, verify them, and promote finished artifacts deliberately.

Build One Predictable Model Path and Cache Policy

Choose one canonical mount such as /srv/models on Linux or a stable drive letter on Windows, and make it available before Ollama, vLLM, LM Studio, or development containers start. Map each tool's model and cache setting explicitly. A symlink is acceptable only when the mount check runs first and the destination never changes between reboots.

Community operators considering a separate NAS repeatedly identify model-load time as the boundary. In one AI workstation and NAS discussion, contributors recommended keeping frequently used models on local NVMe because large weights can take minutes to traverse a slower link.

Use an allowlist for the local hot cache rather than mirroring the entire NAS. After a successful load or copy, verify file size or checksum, then update an atomic alias such as current/model-name. Evict only models that are neither running nor pinned. Keep at least the larger of 15 percent free space or one maximum expected model download so an update cannot fill the boot volume midway.

Protect Manifests and Fine-Tunes, Not Every Download

Back up the information needed to reconstruct the library: source URL or repository ID, exact revision, filename, quantization, checksum, license notes, runtime configuration, and the path used in production. That manifest is small, searchable, and more useful during recovery than a directory full of ambiguously named files.

Back up unique adapters, merged models, calibration data, and evaluation results with normal versioned retention. For public base weights, decide whether recovery time justifies another copy. A slow internet connection or a model that may disappear can make selected weights worth protecting, but mirroring every experiment usually wastes backup capacity.

If the broader AI file layer is still undecided, the ZimaSpace comparison of a personal cloud and local PC storage for AI files is the next planning step. It separates persistent source data and indexes from the machine that performs inference.

Validate Load Time, Offline Behavior, and the Expansion Trigger

Test three paths with the model service stopped: a hot local load, a cold NAS load, and a NAS outage. Record time until the first usable response, peak network throughput, workstation free space before and after, and whether any tool creates a fallback directory on the boot disk. Repeat after a reboot so mount order is tested rather than assumed.

The setup passes when daily models load locally within the expected time, cold models can be staged without manual path edits, unique artifacts restore from backup, and a missing NAS produces a clear failure instead of a silent redownload. Add faster networking or a larger local tier only when measured cold-load delay interrupts work; add NAS capacity when the canonical library approaches its defined free-space floor.

Stop using direct network reads for a workload that repeatedly seeks across model shards, requires predictable low startup latency, or must operate while the NAS is offline. In that case, preserve the NAS as the library and copy complete models into a larger dedicated local SSD before launch.

Final Setup Rule

Keep the canonical model library on shared storage, pin the daily working set on local NVMe, isolate disposable caches, and protect only the artifacts and manifests that cannot be recreated. Expand after measured load time or capacity crosses a written threshold.

NAS & Server Setup

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.