What Your File System Notices
Checksums, journals, snapshots — what ZFS and Btrfs catch that ext4, NTFS and APFS do not.
01How your choice of file system determines whether corruption is caught quietly, or discovered when it's already too late
02Journals Keep the Structure Intact — They Don't Check the Data
Every modern file system on your computer keeps some kind of journal. The idea is straightforward: before committing a write, the file system records its intentions in a separate log. If power dies mid-operation, the journal lets the system replay or roll back that incomplete transaction on the next boot. ext4 on Linux, NTFS on Windows, HFS+ on older macOS — all of them journal metadata this way. It is the reason fsck or chkdsk after an unclean shutdown usually takes seconds rather than hours, and the reason you rarely boot back into a structurally broken volume.
What journalling does not do is check whether the data that made it to disk is the data you actually wrote. A journal records "I am going to write these bytes to this location." It does not ask, afterwards, whether those bytes arrived correctly. If the storage device silently returned corrupted data — a flipped bit, a misdirected write landing in the wrong sector, a sector that reads back differently on retrieval than it reported on write — the journal records none of that. The file system marks the file clean, the operating system reports success, and the corruption is now silently archived in your backup.
This category of failure has a name: silent data corruption, or sometimes a "bit rot" event. It is more common than most people expect. Enterprise storage vendors have tracked it across large fleets for years. It happens on spinning disks. It happens on SSDs. It happens on RAID arrays that dutifully mirror one corrupt copy onto another. A journalling file system is not designed to prevent it; it was never designed to see it.
| Format | Data checksums | Snapshots | Repairs from redundancy |
|---|---|---|---|
| ZFS | Yes — data and metadata | Yes, native | Yes, with a mirror or raidz |
| Btrfs | Yes — data and metadata | Yes, native | Yes, with a redundant profile |
| ext4 | Metadata only | No (LVM below it) | No |
| NTFS | Metadata only | Volume snapshots (VSS) | No |
| APFS | Metadata only | Yes, native | No |
03What ZFS and Btrfs Actually Do Differently
ZFS and Btrfs were designed with a different premise: that hardware lies, and the file system has to verify everything itself. Both do this through end-to-end checksumming. Every data block — and every metadata block — is checksummed when written. When the block is read back, the checksum is recalculated and compared. If they do not match, the file system knows corruption has occurred. This is not a theoretical feature on paper; it is enforced on every read by default.
ZFS, developed at Sun Microsystems and later open-sourced, goes further still. Its architecture is copy-on-write at the block level: rather than modifying a block in place, ZFS writes the new version to a fresh location and updates the pointers only once the write is confirmed. This means a power failure never leaves a partially overwritten block. The old version stays valid until the new version is safely committed. APFS also uses copy-on-write, which is part of why macOS handles sudden shutdown more gracefully than older HFS+ did — but APFS does not add checksums to user data, only to metadata. The structural integrity is protected; the content is not verified.
Btrfs, which arrived later in the mainline Linux kernel, follows a similar copy-on-write model to ZFS and checksums both data and metadata by default. On a single-disk installation this means corruption is detected but, without a second copy of the data, cannot automatically be repaired. Where Btrfs really earns its keep is in RAID configurations or when combined with multiple devices: if it detects a bad block and has a redundant copy, it repairs it silently and logs the event. You can also run btrfs scrub on a mounted volume — a background process that reads every block and validates its checksum — without taking the filesystem offline.
ZFS has the same scrubbing mechanism: zpool scrub. Running it periodically is how administrators find and fix corruption before it propagates. The scrub reads the entire pool, validates every checksum, and reports how many errors it encountered. A healthy pool returns zero errors. A pool with errors tells you exactly which files are affected, whether they could be repaired from redundant copies, and what the underlying device reported. This is a fundamentally different class of information than anything ext4, NTFS, or APFS can provide.
A checksum only tells you a block is wrong. Repairing it needs a second copy — which is why a checksumming filesystem on a single disk is a very good smoke alarm and not a fire brigade.
04Snapshots: What the Others Mean, What ZFS Means
The word "snapshot" is used loosely enough to cause real confusion. Windows Volume Shadow Copy, APFS snapshots, LVM snapshots on Linux — all of these create a point-in-time view of a volume, which is genuinely useful. But they operate at the block or volume level, and they inherit whatever state the file system was in at the moment they were taken. If data was already silently corrupt before the snapshot, the snapshot preserves the corrupt version. A snapshot is not a checksum; it is a frozen copy.
ZFS snapshots are different in one important respect: because ZFS has already checksummed every block, the snapshot's integrity is implicitly guaranteed by the same mechanism. When you clone or send a ZFS snapshot to another pool, the receiving pool validates the checksums on arrival. Corruption introduced in transit is detected. This is why ZFS send/receive is used in serious backup pipelines — not just for its efficiency, but because the data arrives verified.
None of this means that ZFS or Btrfs replace a real backup strategy. They protect against silent corruption within the pool; they do not protect against the pool itself being destroyed, stolen, or accidentally deleted. A scrub that finds and repairs a bad block is doing its job — but if the operator drops a zfs destroy command on the wrong dataset, no amount of checksumming helps. File systems notice corruption. They do not notice mistakes.
The old version stays valid until the new version is safely committed.
05What to Take From This
If you are running ext4, NTFS, or APFS — and most people are — you are not in immediate danger, but you are also flying without instruments on the question of data integrity. These file systems are mature, well-tested, and appropriate for most workloads. They simply cannot tell you whether the bytes on disk match the bytes you wrote. For irreplaceable data — photographs, financial records, source code, anything that cannot be reconstructed — the absence of end-to-end checksumming is a real gap. The right response is not panic; it is regular verification through other means: hash files, periodic restore tests, and a backup structure with genuine retention.
If you have the option of choosing your file system — a NAS you are building, a Linux server, a ZFS-on-Linux or TrueNAS installation — the case for ZFS or Btrfs is strong wherever data integrity matters more than simplicity. The operational overhead is real: you need to run scrubs on a schedule, understand the pool configuration, and not resize things casually. But the payoff is a storage layer that actually tells you when something is wrong, names the file, and — if you have redundancy configured — fixes it while you sleep.
The hierarchy is worth keeping clear: journalling means the structure survives a power failure. Copy-on-write means writes are atomic. Checksumming means the contents are verified. Redundancy means errors can be corrected, not just detected. Only ZFS and Btrfs, of the commonly deployed file systems today, give you all four together. Everything else gives you the first two and asks you to trust the hardware for the rest.
Notes from the bench
- [1]
Checksums find the damage. Redundancy is what lets the filesystem do something about it. ↩