Can You Replace a ZFS Cache Device Without Stopping File Shares?

Eva Wong is the Technical Writer and resident tinkerer at ZimaSpace. A lifelong geek with a passion for homelabs and open-source software, she specializes in translating complex technical concepts into accessible, hands-on guides. Eva believes that self-hosting should be fun, not intimidating. Through her tutorials, she empowers the community to demystify hardware setups, from building their first NAS to mastering Docker containers.

Usually yes: L2ARC is a read cache, so it can be removed and replaced online while file shares continue, provided the device is truly cacheโ€”not log or special allocation.

The decision matters when a cache SSD is failing or being upgraded on a live home NAS. The two competing states are removable L2ARC device and misidentified SLOG or special vdev with different risk. Begin with a saved configuration and disposable data, observe one branch at a time, and stop if the test expands data-loss, permission, or availability risk.

Define the Conditions Behind the Zfs L2Arc Cache Replacement Decision

Record the environment before changing anything: software and firmware versions, device identities, mount or network path, free space, permissions, and the observable symptom. The baseline must preserve enough detail to reproduce a cache SSD is failing or being upgraded on a live home NAS.

The first candidate is removable L2ARC device. The second is misidentified SLOG or special vdev with different risk. The current zpool remove behavior defines the mechanism or command boundary used in the test; it does not replace observation from this specific home server.

Write the acceptance condition and stop condition before running the discriminator. A pass must change the evidence predicted by one branch while leaving unrelated services unchanged; a fail must return the system to the saved state rather than trigger a chain of speculative fixes.

Test the Claim Without Lowering the Original Requirement

Use this discriminator: inspect zpool status and device class, remove the cache device, confirm pool health, then add the replacement. Keep workload, client, path, file set, and timing constant so the result is attributable to the changed variable.

Use L2ARC replacement behavior to select the field that can actually separate the branches, then capture its timestamp, exit status, error text, device or snapshot identity, latency, transferred bytes, permissions, and recovery state. A clean command exit is not enough when identity, durability, or application state is the claim under test.

Repeat the test once after a restart, reconnect, remount, or cold cache when that event is part of the original condition. If the first run is destructive or the environment cannot be restored, stop and reproduce on a disposable copy instead.

zpool status -v
zpool remove pool cache-device
zpool add pool cache replacement-device

Interpret Pass, Fail, and Exception Results

PASS: shares stay available and the pool remains healthy while the new cache warms gradually. Record the exact version, identity, and workload that passed so the conclusion stays conditional rather than becoming a universal claim.

FAIL: the device is log, special, or part of a data vdev, or removal produces errors. A fail does not automatically prove the opposite branch when network, memory, permissions, or source consistency can influence both; isolate those shared dependencies before escalating.

EXCEPTION OR AMBIGUOUS RESULT: stop and protect the pool; do not use cache-device instructions on another vdev class. Preserve logs and do not run repair, prune, destroy, repartition, or recursive ownership commands until a recoverable copy exists.

-15% OFF
Single board computer zimaboard2

Confirm the Decision Under the Original Workload

Apply the action matched to the observed branch, then repeat the original condition rather than a reduced substitute. The decision holds only when shares stay available and the pool remains healthy while the new cache warms gradually across two cycles or the relevant reboot, sleep, interruption, or load transition.

Use the snapshot safety windows to check the nearest dependent workflow, but keep the original trigger unchanged. Unrelated datasets, shares, containers, users, and recovery points must retain their previous access and timing.

The stop boundary is explicit: if the device is log, special, or part of a data vdev, or removal produces errors, return to the last verified configuration, retain the evidence, and escalate to a deeper platform or hardware test only when the branch is repeatable.

After the target result holds, compare it with the storage activity windows so the fix does not move risk into a neighboring service. A successful target test with a new backup, identity, timeout, or availability failure is still a failed change.

FAQ

For ZFS L2ARC cache replacement, the remaining searches usually concern will performance drop after replacement, is an slog the same as a cache device, and should shares be paused anyway. The answers below keep those edge cases separate from the primary decision.

The acceptance boundary does not move: shares stay available and the pool remains healthy while the new cache warms gradually. If a follow-up condition changes the filesystem, identity, network path, or application version, repeat only the discriminator affected by that change.

Stop broadening the experiment when the device is log, special, or part of a data vdev, or removal produces errors. At that point, stop and protect the pool; do not use cache-device instructions on another vdev class; preserve the evidence before escalating to the platform, storage, or hardware owner.

Will performance drop after replacement?

Possibly while the new L2ARC warms; primary ARC and underlying storage continue serving reads.

Is an SLOG the same as a cache device?

No. SLOG participates in synchronous write intent and has different replacement and failure implications.

Should shares be paused anyway?

Not normally for a healthy L2ARC change, but pause heavy work if diagnostics show broader I/O instability.

For ZFS L2ARC cache replacement, the practical answer remains conditional: shares stay available and the pool remains healthy while the new cache warms gradually. When the device is log, special, or part of a data vdev, or removal produces errors, stop and protect the pool; do not use cache-device instructions on another vdev class; a partial success that cannot survive the original workload is not compatibility.

Support & Tips

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.