Skip to content
dbexplore

Question: Write-ahead log

Why is pg_wal filling up?

Answered in the first paragraph. Last updated .

Because segments are being created faster than they are being released, and there are only four reasons a segment is not released: a replication slot still needs it, the archive command has not succeeded on it, a retention setting says to keep it, or checkpoints are not completing so nothing has become releasable yet. Write volume alone does not fill this directory. Retention does.

Eliminating the four in the order that costs least

Start with archiving, because it is the one that fails silently and the one that is easiest to confirm. If a command has begun failing, the files marked ready pile up in the status directory and the archiver statistics record the failure count and the last file it tried. A full backup target, an expired credential and a mistyped path all present identically here.

Slots next. An inactive slot with a growing retention is the classic cause, and it is covered at length in why a replication slot grows.

Then the settings. wal_keep_size retains segments unconditionally, with no consumer needed, and a value inherited from a runbook written for a different cluster will hold a fixed amount forever without anything looking wrong.

Last, checkpoints. Segments only become recyclable at a checkpoint, so a cluster writing hard enough that checkpoints cannot keep pace will hold more than its configured budget suggests. That is not a bug and it is not a retention leak; it is the budget being too small, and the counters that separate a checkpoint on schedule from one forced early are in timed versus requested checkpoints.

What not to do while diagnosing it

Never delete files from this directory by hand. They are the record the server needs to recover, and removing the wrong one converts a disk alert into an unrecoverable cluster. There is a supported command for removing archived segments, and it will refuse the ones still required, which is the entire point of using it.

Also resist raising the log budget as a first move. It buys time on a cluster whose checkpoints are behind, and it buys nothing at all on a cluster with an abandoned slot, because that slot will consume whatever it is given.

Write volume is worth understanding separately, and the accounting for it moved between recent majors: the write-ahead log statistics view in PostgreSQL 14 is where those counters first appeared, and later majors moved part of it elsewhere. For the retention side, replication lag and slot health has the alerting.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.