Skip to content
dbexplore

Glossary: Replication

Replication slot invalidation

Also called: invalidated slot, lost slot.

Definition, revised in place. Last updated .

A replication slot is a server-side bookmark recording how far a consumer has read, and the primary keeps every write-ahead log segment from that point onward so the consumer can resume. Invalidation is what happens when the server decides it will not keep them any longer. The slot’s wal_status becomes lost, the retained segments are removed, and the slot can no longer be used to resume anything: whatever was reading through it must be rebuilt from a fresh copy of the data.

A deliberate failure, chosen over a worse one

Invalidation is a safety valve rather than a fault. A slot with no consumer holds log forever, and the end of that road is a full disk on the primary, which stops writes for every client. max_slot_wal_keep_size puts a ceiling on how much log a slot may pin, and when a slot crosses it the server sacrifices the slot instead of the cluster. The decision is correct and the surprise is entirely in when you learn about it.

Log volume is not the only trigger. A slot can be invalidated because the rows its consumer still needs were removed by vacuum, which is the horizon conflict a standby with feedback disabled runs into; because the objects it depends on were dropped, in the logical case; and, from PostgreSQL 18, because it sat idle longer than idle_replication_slot_timeout allows. Those have nothing in common operationally except the outcome, which is why treating invalidation as one alert with one runbook goes wrong.

What makes it costly is the asymmetry. Getting into this state takes one unattended weekend. Getting out of it means a new base copy of the data, which on a large cluster is hours of transfer and a second set of decisions about when to take it.

Reading a slot before it becomes a problem

Create a slot with nothing attached and the two columns that matter are already visible.

SELECT pg_create_physical_replication_slot('standby_a', true);
SELECT slot_name, active, wal_status,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS wal_held
FROM pg_replication_slots;
 pg_create_physical_replication_slot 
-------------------------------------
 (standby_a,0/242C900)
(1 row)

 slot_name | active | wal_status | wal_held  
-----------+--------+------------+-----------
 standby_a | f      | reserved   | 248 bytes
(1 row)

active is false because nothing is streaming through it, and that combination, an inactive slot with a growing wal_held, is the state to alert on. Waiting for wal_status to reach lost is waiting for the damage. The intermediate value extended means the slot is already past the configured ceiling and surviving only because those segments have not been recycled yet.

From PostgreSQL 17 the view also carries invalidation_reason, which names which of the causes above applied, and inactive_since, which finally answers how long a slot has been abandoned. Before 17 there is no timestamp, and the honest proxy is when the consumer was last seen elsewhere.

Catching it while it is still recoverable

Thresholds that fire early enough to act on, and what to do about a slot that is inactive but legitimate, are covered in replication lag and slot health. If the reason a slot is holding the primary back is a cleanup cutoff rather than log volume, the mechanism is the xmin horizon, and the fix is a different one entirely.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.