Setting up the warnings that tell you a drive is failing — before you find out the hard way on restore day.

01Make the Machine Do the Watching

Most drives announce their own decline well in advance. The information is there — reallocated sectors climbing, raw read error rates spiking, spin-up time creeping upward — but it sits in firmware registers that nobody checks unless something is already wrong. That is exactly backwards. The goal is to get the machine to push a warning to you rather than waiting for the moment a restore fails, or a folder simply isn't there anymore.

The starting point is SMART. Every spinning hard drive and nearly every solid-state drive manufactured in the past two decades exposes Self-Monitoring, Analysis and Reporting Technology data over the ATA or NVMe interface. A handful of those attributes carry genuine predictive weight — pending and reallocated sector counts in particular — and those are the ones worth setting a threshold on. Running a SMART query once manually is not monitoring; it tells you the state of one moment and nothing more. What you want is a daemon or scheduled task that queries those attributes on a regular cycle and sends an alert when any watched value moves.

On Linux, smartd does exactly this. Shipped with the smartmontools package, it runs as a system service, polls every drive at a configurable interval, and can email an administrator when a threshold is crossed or when a drive itself raises an internal flag. The default configuration is conservative; the documentation walks through per-device directives if you need tighter control. On Windows, the Task Scheduler can run smartctl on a schedule and pipe the output to a log, though a purpose-built service makes alerting easier. macOS ships a background process that reads SMART status and surfaces a basic warning in Disk Utility, but it exposes only a binary pass/fail; for anything more granular, smartmontools installs cleanly via Homebrew.

The point of all of this is not to automate a recovery — it is simply to know. A rising reallocated sector count does not mean the drive dies tomorrow; it means the drive has started a process that ends in failure, and you now have a window.

02Logs, Scrubs and the Filesystem Layer

SMART watches the drive. But some failures happen above the drive — at the filesystem level — and require a different kind of watching. Bit rot, the silent corruption of data that has not been read in a long time, produces no SMART event because the drive's electronics never knew anything was wrong. The data was written correctly, stored correctly, and then the magnetic domain or flash cell changed state on its own.

ZFS and Btrfs both keep checksums of every block they write. Running a scrub periodically — ZFS has zpool scrub, Btrfs has btrfs scrub — reads every block, verifies it against its stored checksum, and reports mismatches. On a pool with redundancy, ZFS will repair the corrupt block silently from the redundant copy. Without redundancy it still reports the corruption, which is infinitely better than silence. Scheduling a monthly scrub and routing its output to a log you actually read turns the filesystem into an active participant in your health monitoring. On ext4, NTFS, and APFS, what the filesystem notices is considerably less — there is no equivalent to a checksum scrub, so SMART monitoring becomes the primary early-warning layer.

Beyond SMART and scrubs, the operating system's own logs are worth checking. On Linux, dmesg and the systemd journal both surface ATA error messages — input/output errors, command timeouts, reset events — that often precede a visible SMART change. Setting up log monitoring to watch for the strings ata, blk_update_request, or I/O error and alerting on them costs very little effort and catches a class of problem that SMART misses.

ZFS and Btrfs both keep checksums of every block they write.

03Turning Alerts into a Habit

The final piece is the simplest and the most neglected: the alert has to go somewhere you will see it. An email to a mailbox that nobody reads is not monitoring; it is documentation for a post-mortem. Route SMART alerts to a channel you check daily — a real inbox, a phone notification, a home-automation alert if that is your preference. For a small office, a daily digest of SMART status and the result of any overnight scrub gives an administrator a two-minute health check that catches most problems well before they become recoverable-or-not decisions.

None of this replaces a working backup. A drive can fail catastrophically without warning, and SMART has a documented false-negative rate — drives do die clean. But monitoring closes the gap between "the drive was already showing signs" and "nobody was looking." The goal is to make sure that when a drive starts its slow decline, you are the first to know — not the last.

The test of a monitoring setup is not whether it can show you a dashboard. It is whether it woke somebody up the last time a counter moved.