Replication
Replication lag and slot health
For anyone whose standby is behind and who needs to know whether that is a network problem, a replay problem or a disk problem.
Byte lag, time lag and slot retention are three different measurements. Which to alert on, how to read them, and why a slot can end a cluster.
Reference page, revised in place. Last updated .
Lag is four numbers, not one
Write-ahead log records leave the primary, cross the network, land on the standby’s disk, and are then replayed into its data files. A record can be stuck at any of those four points, and “replication lag” as a single number hides which.
PostgreSQL exposes the stages separately. On the primary, pg_stat_replication has one row per connected walsender carrying sent_lsn, write_lsn, flush_lsn and replay_lsn, which are the log positions the standby has respectively been sent, written to its operating system, flushed to durable storage, and applied. Alongside them sit write_lag, flush_lag and replay_lag, which are intervals rather than byte counts: they measure how long it took the standby to confirm each stage for a recently written record.
The distinction matters because the three failure modes look identical on a dashboard that only plots one line. A gap between sent_lsn and write_lsn is the network or a saturated standby disk. A gap between flush_lsn and replay_lsn is single-threaded recovery falling behind, or recovery deliberately paused, or a replication conflict. And a standby that is perfectly caught up on a primary that is not writing anything will show a large time lag for no reason at all, which is the single most common false alarm in Postgres monitoring.
What to run on the primary
SELECT application_name,
client_addr,
state,
sync_state,
pg_wal_lsn_diff(pg_current_wal_lsn(), sent_lsn) AS pending_send_bytes,
pg_wal_lsn_diff(sent_lsn, flush_lsn) AS unflushed_bytes,
pg_wal_lsn_diff(flush_lsn, replay_lsn) AS unreplayed_bytes,
write_lag, flush_lag, replay_lag
FROM pg_stat_replication
ORDER BY pending_send_bytes DESC NULLS LAST;
pg_wal_lsn_diff() returns the byte distance between two log positions as a numeric. pg_current_wal_lsn() is a primary-only function and raises an error during recovery, so this query belongs on the writer and nowhere else.
state is worth reading before the numbers. A walsender in catchup is still feeding a standby that fell behind and is working through the backlog; one in streaming is keeping up. sync_state tells you whether a commit on the primary is waiting for this standby, and a synchronous standby with a growing flush_lag is not a monitoring problem, it is an availability problem happening right now.
What to run on the standby
The standby cannot see the primary’s current position, so it measures differently:
SELECT pg_is_in_recovery() AS in_recovery,
pg_last_wal_receive_lsn() AS received,
pg_last_wal_replay_lsn() AS replayed,
pg_wal_lsn_diff(pg_last_wal_receive_lsn(),
pg_last_wal_replay_lsn()) AS replay_backlog_bytes,
pg_last_xact_replay_timestamp() AS last_replayed_commit,
now() - pg_last_xact_replay_timestamp() AS apparent_time_lag,
pg_get_wal_replay_pause_state() AS pause_state;
apparent_time_lag is the number most dashboards show and the one to treat with suspicion. It is the age of the last commit that was replayed, so on an idle primary it grows by one second per second while the standby is entirely healthy. Pair it with replay_backlog_bytes: if the byte backlog is zero and the time lag is minutes, the primary is quiet and nothing is wrong. If both are large, you have real lag.
pause_state catches the case nobody expects, where recovery was paused by hand or by recovery_min_apply_delay and never resumed. It returns not paused, pause requested or paused, and it only runs during recovery.
Slots are a different problem
A replication slot is a promise. It tells the primary to keep write-ahead log segments, and to hold back the cleanup horizon for logical slots, until the consumer confirms it has them. That promise is kept even when the consumer has been gone for a week, which is why an abandoned slot is the most reliable way to fill a Postgres data directory.
SELECT slot_name,
slot_type,
database,
active,
active_pid,
wal_status,
pg_size_pretty(safe_wal_size) AS headroom,
inactive_since,
conflicting,
invalidation_reason,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(),
restart_lsn)) AS retained_wal
FROM pg_replication_slots
ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC;
That is another primary-side query, for the same reason as the first one. Read wal_status first. reserved means the log the slot needs is within max_wal_size. extended means the log is being kept beyond that because of this slot, which is the early warning. unreserved means the slot is now past max_slot_wal_keep_size and the segments it wants are on borrowed time. lost means they are gone and the slot can never be resumed; the consumer must be rebuilt from a fresh copy.
safe_wal_size is the headroom in bytes before that happens, and it is the right thing to alert on, because it counts down toward zero regardless of how large the cluster is. inactive_since and invalidation_reason both arrived in PostgreSQL 17, and version 18 adds idle_replication_slot_timeout so that a slot nobody has read from for long enough is invalidated deliberately rather than by running out of disk. On 16 and earlier, active plus restart_lsn is all you get, and a slot with active false and an old restart_lsn is the one to investigate.
Those two columns are the difference between reading a slot’s obituary and searching the log for it, and what each invalidation reason actually tells you to do is worth having straight before one fires.
max_slot_wal_keep_size is the setting that turns a cluster-ending problem into a broken replica. Left unset, a slot can consume the entire filesystem and stop the primary. Set to a size you can afford, the primary survives and the slot is invalidated instead. That is a real trade-off and it should be made deliberately, at a size somebody chose rather than at the default of no limit at all.
Logical slots and the cleanup horizon
Physical slots retain log segments. Logical slots retain log segments and hold back catalog_xmin, which stops vacuum from removing catalog rows a decoder might still need. A logical slot whose consumer has stalled therefore causes bloat at the same time as it causes disk growth, and the bloat is on the catalog, which affects planning for every query in the database.
For logical replication specifically, pg_stat_replication_slots reports how much decoding work is spilling to disk or being streamed early, which is the difference between a slow consumer and a slot that is generating large transactions:
SELECT slot_name, spill_txns, spill_count,
pg_size_pretty(spill_bytes) AS spilled,
stream_txns, stream_count,
pg_size_pretty(stream_bytes) AS streamed,
stats_reset
FROM pg_stat_replication_slots
ORDER BY spill_bytes DESC;
Rising spill_bytes means transactions are exceeding logical_decoding_work_mem and being written to temporary files before they can be sent. Raising that setting is usually cheaper than the I/O the spilling costs.
Conflicts, feedback, and the choice between them
A standby serving read queries has a conflict with recovery whenever the primary removes a row version that a running query on the standby still needs. PostgreSQL resolves that by waiting max_standby_streaming_delay and then cancelling the query. The alternative is hot_standby_feedback, which sends the standby’s oldest snapshot back to the primary so the row is never removed in the first place.
That is a genuine trade, and both directions have a cost. Feedback off means queries on the standby get cancelled. Feedback on means a long query on the standby pins the primary’s cleanup horizon, and you get bloat and rising transaction age on the writer because of a report running on a replica. pg_stat_database_conflicts counts the cancellations by cause, and the backend_xmin column of pg_stat_replication shows what the feedback is currently costing you. Pick the side that matches what the standby is for, and monitor the side effect of whichever you picked.