Why Does Immich Metadata Grow During Family Photo Backup?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

Immich metadata grows during family photo backup because each original creates application records and may also create thumbnails, search vectors, face data, and other derived state.

That growth is not one uniform percentage of the original library. A household with many small images, long videos, numerous faces, or extensive search processing can produce a different overhead profile from another family with the same source terabytes. Separate database state from generated files before deciding whether growth is expected.

Every Asset Adds Durable Application Records

The database needs records that connect an asset to its owner, path, timestamps, albums, permissions, and other application-visible properties. As the number of assets and relationships grows, so does this durable state even when the original files are stored elsewhere.

A storage-sizing overview that separates database overhead from thumbnails and originals is useful because these roles scale differently. Its example figures should be treated as deployment observations rather than a guaranteed ratio for another family library.

Asset count is therefore a better starting variable than source gigabytes for some metadata questions. Ten thousand large videos and ten thousand small photos may use very different original capacity, yet both still require asset-level records and relationships in the application database.

Generated Browsing Files Add a Separate Storage Curve

Timeline browsing depends on smaller representations that are faster to display than opening every original. These generated files are not database metadata in the narrow sense, but they are often perceived as โ€œImmich overheadโ€ because they grow alongside the library and are managed by the application.

The four-service storage layout in a home-lab deployment distinguishes photos, generated media, and database placement. That separation is operationally important because high-churn derivative files and critical database state do not have the same backup or performance role.

Do not estimate this curve from original bytes alone. Thumbnail count follows assets and enabled sizes, while encoded-video output follows video compatibility and transcode settings. Measure each generated directory separately after the same cohort has completed processing.

Search and Face Features Add Index State

Semantic search and facial features create numerical representations and relationships that make visual content discoverable without rewriting the original image. More processed assets, detected faces, and enabled analysis features therefore add database and model-related state over time.

The explanation of semantic embeddings shows why a visual search index can grow even when filenames and folders stay unchanged. The model transforms each eligible image into a reusable representation, which is then available for later comparison with text queries.

This state should not be confused with a second full-resolution copy. If search-related database growth continues rapidly after asset count, feature settings, and model inventory remain stable, investigate maintenance, duplicated processing, or another database mechanism instead of assuming normal indexing explains it.

Family Organization Adds Relationships, Not Just Files

Albums, people names, sharing relationships, favorites, edits, and other user actions can expand application metadata independently of new originals. Two families with identical media can therefore have different database footprints because one uses more organizational and sharing features.

ZimaSpaceโ€™s family-photo overview of photo organization layers highlights the distinction between centralized originals and the searchable people, places, events, and albums layered above them. Those relationships are part of the user experience and must be considered in recovery planning.

The mechanism stops explaining large unexplained growth in logs, container writable layers, temporary files, or duplicate source assets. Those categories have different causes and should be measured outside the database-and-derivatives model rather than folded into a single โ€œmetadataโ€ number.

Measure Growth by Storage Role

Take a baseline before a representative import: original media bytes and count, database size, thumbnail or preview storage, encoded-video storage, model cache, backups, and temporary or log space. Repeat after the same cohort has completed its enabled background jobs and again after normal household use.

A ZimaSpace family backup workflow reinforces why originals and essential application state must be protected together while regenerable outputs can be treated differently. Storage accounting should follow recovery value as well as byte count.

Accept growth when the change can be reconciled to new assets, derivatives, database records, and enabled features. Investigate further when one role grows without matching asset or feature activity, or when the measured total diverges materially from the sum of the known storage roles.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.