Why Does a Btrfs Balance Stall After Moving Data to a Larger Drive?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

A Btrfs balance can appear stalled after a larger drive is added because relocation still needs free chunk workspace and may be moving the entire filesystem.

Replacing or copying data to a larger device does not automatically make every Btrfs block group compact, evenly distributed, or eligible for relocation. A balance works at the block-group level, creates temporary workspace, updates metadata, and can be limited by the slowest device, snapshots, checksums, or another exclusive operation. The first task is to distinguish a genuinely stuck relocation from a slow full balance that is still making progress.

Confirm Whether the Balance Is Running, Paused, or Waiting

Check the balance status, processed chunk count, kernel log, disk throughput, and per-device latency. Record whether the command is running in the foreground, background, paused, cancelled, or automatically resumed after reboot.

A balance may spend a long time relocating one heavily used block group before the visible counter changes. ArchWiki’s Btrfs guidance shows the use of balance status and filesystem usage to distinguish continued work from a command that has already stopped.

If there is no I/O, no status change, and the kernel log reports an error, treat it as stopped rather than slow. Preserve the first error message before restarting the balance with different filters.

Verify That the Larger Device and Filesystem Were Resized

Compare the physical device size, partition size, Btrfs device size, and filesystem allocation. A larger replacement disk may still expose the old partition boundary or old Btrfs device size.

Confirm each layer in order: hardware capacity, partition table, block device, Btrfs device inventory, and filesystem allocation. SUSE documents that the device must be enlarged before the filesystem, so a balance cannot use capacity that Btrfs does not yet see.

The ZimaSpace article on capacity not growing after drive replacement provides the adjacent layered check before any relocation is blamed.

Distinguish Free Bytes From Free Chunk Workspace

Compare total device size with the amount already assigned to Btrfs block groups. A filesystem can show free bytes to users yet lack a completely unallocated region large enough to create the temporary block group needed for relocation.

The Btrfs balance documentation explains that relocation needs completely unused block-group workspace; this differs from ordinary file-level free space and can produce ENOSPC during balance.

If workspace is constrained, first reclaim completely unused block groups with a narrow usage filter rather than starting another full balance. Do not delete snapshots indiscriminately until their contribution to chunk allocation is measured.

-15% OFF
Single board computer zimaboard2

Check Whether a Full Balance Was Started Without Filters

Review the original command. A balance without data or metadata filters attempts to relocate the entire filesystem, even when the real goal is only to compact lightly used chunks or move allocations away from one device.

A full balance can take many hours or days because every selected block group is rewritten. The Linux btrfs-balance manual warns that running without filters moves data and metadata across the whole filesystem and updates all block pointers.

Use the status output and previous shell history to identify the active filters. Do not cancel and restart repeatedly, because interrupted balances can leave partially filled block groups that continue consuming workspace.

Check Slow Devices, Errors, and Competing Exclusive Operations

Inspect SMART data, transport errors, link resets, USB or SATA timeouts, and per-device latency. Balance speed is constrained by reads from old locations, writes to new locations, checksum verification, and metadata updates.

Also check for scrub, device add or remove, filesystem resize, snapshot deletion, send or receive, and other storage work. Red Hat’s Btrfs administration guide describes device changes and balance as relocation operations, so overlapping maintenance can create heavy contention even when no disk has failed.

If one device shows repeated resets or extreme latency, pause the balance and diagnose that path before forcing more relocation. Continuing through an unstable device can turn a performance problem into a recovery problem.

Use Narrow Filters and Limits for a Controlled Restart

After preserving the current state, start with the lowest-risk filter that targets empty or lightly used block groups. Limit the number of chunks per run so each result can be observed before increasing scope.

Increase the usage threshold gradually and treat data and metadata separately. Metadata relocation can generate many additional updates and does not need to be compacted aggressively merely because the average utilization looks low.

Pause, resume, or cancel through the supported balance controls rather than killing the process. Confirm that the current block group completes and that the saved balance state matches the next action.

Verify Distribution and Capacity After the Balance

Compare per-device allocation, data and metadata profiles, unallocated workspace, filesystem usage, and the number of relocated chunks before and after the controlled balance.

A successful outcome is not necessarily perfectly equal byte usage across every drive. The Linux kernel overview lists integrated multi-device support and online resize, so the final test should focus on valid profiles, usable allocation, healthy devices, and enough workspace for future writes.

The issue is resolved when the balance completes or reaches the intended filtered scope, the larger device receives new allocations, free chunk workspace is restored, and normal writes, snapshots, scrub, and reboot work without the balance restarting unexpectedly.

Support & Tips

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.