Skip to content
dbexplore

Question: Write-ahead log

How do I know if WAL archiving is working?

Answered in the first paragraph. Last updated .

Read the archiver statistics view. It carries the number of files archived, the name and time of the last success, and the matching pair for failures. A healthy cluster shows a recent last-success time and a failure count that is not moving. Both halves matter, and the second half alone is not enough: an archiver that has stopped attempting anything has no failures either.

What the counters catch, and the one they miss

An expired credential, a full destination or a mistyped path all present the same way: the failure count climbs, the last-failed name and time move, and files marked ready accumulate in the status directory beside the log. That is the easy case and any alert on the failure count finds it.

The documented gap is more interesting. If the archive command is killed by a signal, or the shell cannot find it at all, the archiver process itself aborts and is restarted, and that failure is not recorded in the statistics. So a cluster whose archive command was renamed by a package upgrade can show zero failures forever while archiving nothing. This is the reason the alert has to be on the age of the last successful archive, which keeps rising in every failure mode including that one.

There is a second assumption worth dropping. The documentation warns that files are not guaranteed to be archived in order, particularly after a promotion or a crash, so you cannot infer that everything older than the last archived file is safely away.

Proving it rather than believing it

The command has to be honest. It must return failure when it fails, and it must refuse to overwrite a file that is already in the archive, because that is what protects you from two servers writing into the same destination. A command that returns success on failure produces a gap in the archive that nothing will report until a restore needs it.

Which is the real test. Archiving that has never been restored from is a belief, not a backup: take a restore to a scratch host on a schedule and measure how long it takes, because that number is your actual recovery time.

PostgreSQL 15 added a module interface for archiving and gave the archiver its own wait events, so what the archiver is waiting on is now visible without guessing; the statistics view it reports through is unchanged, which is why the alerting above still applies as written.

When archiving does stall, the segments stay behind and the disk is the next thing to go, which is why pg_wal fills up. Replication lag and slot health covers alerting on retention as a whole.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.