Btrfs Metadata Recovery Checklist for a Nearly Full Home Server

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

The safe approach is to treat stabilize writes, recover workspace, run only filtered balance, and verify data before normal service as a sequence of observable gates, not a single command.

On a nearly full Btrfs home-server filesystem, the practical risk is Btrfs reports ENOSPC or becomes read-only while metadata space is exhausted. Record the current identity and recovery point, start with the least invasive discriminator, interpret pass and fail results before changing another variable, and stop when storage becomes unstable or the only recoverable copy would be exposed. The workflow below ends only after the original workload succeeds or the evidence reaches an escalation boundary.

Stabilize the filesystem before attempting repair

Stop containers, downloads, snapshots, and log-heavy jobs that write to the affected filesystem. Save kernel errors and btrfs device stats output somewhere else. If the filesystem remounted read-only or reports checksum, parent transid, or I/O errors, keep it read-only until you have a recoverable copy.

Do not begin with btrfs check --repair, a full balance, defragmentation, or mass deletion. The immediate question is whether valid filesystem state lacks allocation workspace or whether storage errors are damaging metadata; repair-first action can make that distinction harder and consume the remaining space.

The safety gate passes when write-heavy services are stopped, important data has another copy, and you know which block device and mountpoint you are examining. Escalate to recovery imaging if the device resets, disappears, or accumulates read errors.

Read allocation rather than the headline free-space number

Run btrfs filesystem usage -T /mount, btrfs filesystem df /mount, btrfs device usage /mount, and inspect recent kernel messages. Compare allocated versus used metadata and the unallocated space available on every device; ordinary df alone cannot show whether Btrfs can allocate another metadata chunk.

A targeted balance needs completely unused workspace. A detailed targeted Btrfs balance guide explains that an unfiltered balance rewrites all eligible block groups and that the goal is to maintain device-level unallocated space, not merely delete a large file and assume metadata can grow.

If metadata is high but unallocated space remains, a small filtered balance may reclaim empty or lightly used chunks. If no device has workspace, first remove safely expendable data or snapshots in small batches, or add a temporary device appropriate to the filesystem profile; do not launch relocation that cannot finish.

Recover workspace with the least invasive action

Start with deletion of disposable files that are not retained by snapshots, then delete only confirmed unneeded snapshots. Sync and recheck usage after each small change. If a balance is justified, begin with btrfs balance start -dusage=0 -musage=0 /mount or another narrow filter chosen from the observed allocation, not a full balance.

The Linux manual description of filtered balance behavior notes that filters limit relocation and that ENOSPC can occur when balance itself lacks workspace. Watch btrfs balance status and kernel logs. If relocation increases errors, stalls with device faults, or consumes the last safety margin, cancel it and return to read-only recovery.

Do not stack filters and deletions without measuring between steps. The recovery branch succeeds when metadata has headroom, unallocated space exists on the necessary devices, and a small write completes without a new ENOSPC or forced read-only event.

Verify data and prevent an immediate relapse

Restart only one low-risk service and reproduce the workload that originally filled metadata, such as snapshot creation or many small file changes. Recheck usage and kernel logs after the workload and after a reboot. A mount that works once but returns to read-only under normal churn is not recovered.

Use the ZimaSpace method for separating snapshots or live files use NAS space before changing retention. Snapshot-held extents can make deletion appear ineffective, while live small-file churn can keep metadata pressure high; the correct policy depends on which state the measurements prove.

Resume normal service only after the filesystem remains writable, device stats stop increasing, a representative file restores or hashes correctly, and monitoring alerts before the same margin disappears. Escalate persistent structural errors to a Btrfs recovery specialist and work from a clone rather than repeating repair commands on the only copy.

Support & Tips

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.