PostgreSQL 15
The archiver says what it is waiting on
For anyone whose archive command is a shell script and whose only evidence about it is a counter that has stopped rising.
Reference page, revised in place. Last updated .
Two counters and a shrug
Archiving is the part of a PostgreSQL deployment that everybody configures once and then depends on for years. A segment of write-ahead log fills, the archiver hands it to whatever you nominated, and if that works the segment can be recycled. If it does not work, segments accumulate, and the cluster eventually runs out of disk.
The instrument for this has been the same for a long time: a count of segments archived, a count of failures, and the name and time of the most recent of each. Those four facts answer the question “is archiving working” and answer nothing else. In particular, when the count stops rising, they do not distinguish between an archiver waiting on a command that has hung, an archiver that has nothing to do, and an archiver that is failing so fast the failure count is the thing moving.
PostgreSQL 15 improved this in two directions at once and left the statistics view alone. The Streaming Replication and Recovery section of the PostgreSQL 15 release notes records the library interface, and the Monitoring section records wait events for the local shell commands the server runs, of which archiving is the one that matters most.
The configuration half
SHOW archive_library;
ERROR: unrecognized configuration parameter "archive_library"
SHOW archive_library;
SELECT name, setting, context FROM pg_settings WHERE name IN ('archive_mode', 'archive_command', 'archive_library')
ORDER BY name;
archive_library
-----------------
(1 row)
name | setting | context
-----------------+---------+------------
archive_command | sleep 4 | sighup
archive_library | | sighup
archive_mode | on | postmaster
(3 rows)
Empty means the shell command is in use, which keeps every existing deployment working untouched. The alternative is naming a loadable library that the server calls directly instead of forking a shell for every segment.
Why that matters is a question of failure modes rather than speed. A shell command that is handed a segment has no way to tell the server anything except an exit status. It cannot say that it succeeded but only after retrying, it cannot say that it is going to be slow, and a command that hangs takes the archiver with it. A library runs in the archiver process, can keep state between calls, and can report a failure in a way the server understands.
It is also the sharper edge of the two, because a library that misbehaves misbehaves inside the server process rather than in a child of it. That is the trade, and it is why the shell command remains the default and remains supported.
The wait-event half, which is where the diagnosis is
This is the change that shows up during an incident. The fixture below is a server whose archive command takes several seconds to do nothing, which is a faithful model of an archive command that is waiting on a network filesystem or an object store.
SELECT pg_switch_wal();
SELECT pg_sleep(2);
SELECT backend_type, state, wait_event_type, wait_event
FROM pg_stat_activity WHERE backend_type = 'archiver';
On 14:
pg_switch_wal
---------------
0/2000958
(1 row)
pg_sleep
----------
(1 row)
backend_type | state | wait_event_type | wait_event
--------------+-------+-----------------+------------
archiver | | |
(1 row)
On 15:
pg_switch_wal
---------------
0/2424028
(1 row)
pg_sleep
----------
(1 row)
backend_type | state | wait_event_type | wait_event
--------------+-------+-----------------+----------------
archiver | | IPC | ArchiveCommand
(1 row)
The archiver is present in the activity view on both versions, and on 14 it has nothing to say about itself. No state, no wait event, no indication that it is inside a shell command that has been running for seconds. On 15 the same row names the wait.
That single field changes the shape of the investigation. On 14, an operator who sees the archived count flat has to go outside the database to find out why: look for the archive process in the operating system’s process list, check whether the mount it writes to is responding, read the log. On 15 the database itself distinguishes an archiver blocked in its command from an archiver with nothing to do, and it does so through the same view every other stuck process is diagnosed from.
The same treatment was given to the commands used during recovery and at the end of it, which are the other places a shell command can quietly hang. A standby that appears stuck partway through recovery is a much less mysterious object when the view says it is inside the restore command.
The view that did not change
It is worth confirming rather than assuming, because a release that improves archiving in two ways invites the assumption that the reporting improved too.
SELECT string_agg(attname, ', ' ORDER BY attnum) AS columns
FROM pg_attribute WHERE attrelid = 'pg_stat_archiver'::regclass AND attnum > 0;
On 14:
columns
---------------------------------------------------------------------------------------------------------------------
archived_count, last_archived_wal, last_archived_time, failed_count, last_failed_wal, last_failed_time, stats_reset
(1 row)
On 15:
columns
---------------------------------------------------------------------------------------------------------------------
archived_count, last_archived_wal, last_archived_time, failed_count, last_failed_wal, last_failed_time, stats_reset
(1 row)
Identical. Which has a consequence worth stating plainly: nothing in the statistics view tells you whether a cluster is archiving through a shell command or through a library. If you are rolling out the library interface across a fleet, the settings are how you confirm the rollout, not the counters.
The counters also still say nothing about the backlog. The number of segments waiting to be archived is not in the view; it is on the filesystem, in the directory of ready markers. A check that watches the archived count and the failure count will not notice a cluster that is archiving successfully but more slowly than it is generating, which is the failure that ends with a full disk while every counter looks healthy.
What a stuck archiver does to everything else
The reason a wait event on this process is worth more than it looks is that an archiver which stops making progress does not stay a local problem for long.
Segments that have not been archived cannot be recycled, so the write-ahead log directory grows. On a cluster with generous storage that is slow and survivable. On a cluster sized for a steady state it is a countdown, and the end of it is a server that cannot write at all, because a database that cannot create a new segment cannot commit.
The countdown has no natural alarm attached to it. The archived count is flat, which is the same thing it does overnight on a quiet system. The failure count is not rising, because a command that hangs has not failed. Disk usage climbs, which is monitored on most fleets, and by the time it is the thing that alerts there is often less time left than anybody expects, because the rate of growth is the rate the workload generates log rather than anything related to archiving.
There is a second-order effect that catches teams whose retention is driven by archiving. If the archive is what the backup depends on, then an archiver that has been stuck for six hours has left a six-hour hole in the recovery window, and nothing in the backup tooling knows that. A restore test run the next morning would find it. Almost nobody runs one the next morning.
This is the whole argument for the wait event. It turns a condition whose only symptom is the absence of a symptom into something a query can see, in the same view the rest of the cluster is diagnosed from, and it does so at the moment the archiver enters the command rather than after a timeout somebody had to choose.
The backlog nobody can query
Even with the wait event, one number is still outside the database, and it is the number that says how bad the situation is.
The server keeps a marker for each segment that is ready to be archived, and the count of those markers is the backlog. It lives on the filesystem, not in a view, and there is no supported query that returns it. A cluster archiving successfully but more slowly than it generates log has a backlog that grows steadily while every counter in the statistics view looks healthy, and that is the failure this page’s improvements still do not cover.
Getting it requires reading a directory from outside the database, which most monitoring agents can do, and comparing it against the rate at which segments are being produced. Neither number is difficult. The reason it is so often missing is that the statistics view looks like it ought to cover this and does not.
What to actually monitor
Four signals, and only one of them is new in this release.
- The failure count as a rate. Any sustained rise is an outage in slow motion.
- The age of the last archived segment. A flat count is only meaningful against how much log the cluster is producing.
- The archiver’s wait event, on 15 and later, sampled often enough to catch a command that hangs rather than fails.
- The backlog of segments not yet archived, which has to come from outside the database.
A change of archiving mechanism is a good moment to check all four, because the library interface removes the fork per segment and can therefore be substantially faster, and a rollout that fixes throughput will also hide any backlog that existed before it.
What it costs
Nothing to read. The wait event is recorded in the same shared memory every other process reports through, and asking for it is an ordinary query against the activity view.
The library interface costs whatever the library costs, and it moves that cost inside the server. The shell command costs a process fork per segment, which on a cluster generating segments rapidly is not negligible, and it is the reason the alternative exists at all. Neither choice changes what the statistics view reports, and for a fleet the practical consequence is that the mechanism belongs in your inventory of what each cluster is, alongside the log settings the same release changed the defaults of.