Why Does NAS Throughput Fall When an AI Indexer Scans Millions of Small Files?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

NAS throughput falls during small-file indexing because fixed metadata and open-close costs dominate before storage can reach efficient sequential transfer rates.

An indexer scanning one terabyte in a few large archives can stream data, but the same bytes spread across millions of documents require millions of lookups. Each path may trigger directory traversal, permission checks, attributes, opens, tiny reads, closes, hashing, and index writes. The workload becomes operations-per-second and latency bound rather than purely bandwidth bound.

Each File Adds Work That Does Not Scale With Its Size

Before reading content, the client and NAS resolve a path, inspect metadata, enforce permissions, and open a handle. These fixed operations cost nearly as much for a two-kilobyte note as for a large video, so useful bytes per request collapse as average file size shrinks.

The TableFS work on metadata-dominated workloads was designed for workloads dominated by metadata and tiny files, demonstrating that conventional local filesystems can bottleneck on namespace operations even when the underlying storage can transfer data much faster.

A network filesystem adds protocol exchanges and server-side locking or cache validation. Parallelism can hide some latency, but too many workers deepen queues, evict useful metadata, and make interactive NAS requests wait behind bulk enumeration.

Large Directories and Random Access Break Sequential Efficiency

Millions of entries expand directory indexes and inode working sets beyond cache. The scan jumps among metadata blocks and small extents, reducing read-ahead effectiveness and forcing HDD seeks or scattered SSD requests instead of long sequential transfers.

Research on scalable file directories examines directories containing millions to billions of small files and distributes metadata growth across partitions. Its design illustrates that namespace scalability is a separate problem from raw device bandwidth. This distinction remains visible during later household testing.

The AI pipeline adds another random stream when it writes hashes, OCR text, thumbnails, or vector database records. Read and write queues compete, and synchronous database commits can pause ingestion even while the disks show unused peak sequential throughput.

Cache Churn Spreads the Slowdown to Other NAS Users

Directory entries, attributes, file data, model pages, and index buffers all compete for RAM. A broad scan can replace hot household files in cache, while antivirus, thumbnailing, or checksumming duplicates reads triggered by the same new access events.

An analysis of small-file overhead explains how large populations of sub-64KB objects create metadata and request overhead that bulk throughput metrics hide. Grouping work changes the ratio between useful payload and per-object processing. The intermediate result must remain inspectable before automation follows.

The failure boundary is assuming file count alone sets performance. A warm SSD metadata cache, packed archive, local index database, and batched protocol can handle more files than a cold HDD over high-latency SMB. Measure operations and queueing under the actual layout.

-15% OFF
Single board computer zimaboard2

Benchmark Files per Second Separately From Megabytes per Second

Create datasets with equal total bytes but median files of 4KB, 64KB, 1MB, and 64MB, plus shallow and deeply nested directories. Record files per second, metadata operations, network round trips, IOPS, queue latency, cache misses, index writes, and interactive NAS latency.

Use bottleneck separation as the bottleneck framework while repeating with one, four, and sixteen indexer workers, then with content batching and the index database on another device. Keep the file set and cache state explicit.

Choose concurrency at the point where files per second stops improving or interactive p95 latency crosses its limit. If metadata is dominant, reduce repeated stats and batch work; if content reads dominate, optimize storage layout rather than treating sequential bandwidth as the missing capacity.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.