Skip to content
dbexplore

Glossary: Replication

A replication slot is a retention promise

Also called: physical slot, logical slot, pg_replication_slots.

Definition, revised in place. Last updated .

A replication slot is a named marker on the sending server recording how far one consumer has read. From that position onward the server keeps every log segment, and for a slot used by logical decoding it also keeps the catalog row versions needed to interpret them. The marker is durable: it survives a restart, it exists while nothing is attached to it, and it moves only when a consumer confirms progress or somebody drops it.

The guarantee and the liability are the same mechanism

What a slot buys is that a standby which loses its network for an hour can resume rather than needing a new base backup, because the primary held what it needed instead of guessing at a retention window. That is a large operational improvement over sizing a fixed reserve and hoping.

What it costs is that the promise has no expiry. A slot left behind by a standby that was decommissioned, or by a consumer whose process died, goes on reserving log at exactly the same priority as one serving a live replica. The server cannot tell the difference, because from its side there is none.

Logical slots carry a second liability that is easier to miss. As well as pinning log, they pin the oldest transaction id whose catalog versions must remain readable, which holds back cleanup across the whole cluster. A stalled logical slot therefore shows up first as unexplained bloat on tables that have nothing to do with the replicated ones.

Creating one and reading what it is holding

A physical slot with no consumer, which is exactly the state that causes trouble.

SELECT slot_name, lsn IS NOT NULL AS has_position
FROM pg_create_physical_replication_slot('reporting_standby', true);
SELECT slot_name, slot_type, active,
       pg_size_pretty(pg_current_wal_lsn() - restart_lsn) AS wal_held
FROM pg_replication_slots;
SELECT pg_drop_replication_slot('reporting_standby');
     slot_name     | has_position 
-------------------+--------------
 reporting_standby | t
(1 row)

     slot_name     | slot_type | active | wal_held 
-------------------+-----------+--------+----------
 reporting_standby | physical  | f      | 88 MB
(1 row)

 pg_drop_replication_slot 
--------------------------
 
(1 row)

Read the active column carefully, because it is the one that gets alerted on and it means only whether something is connected right now. A healthy standby that reconnects every few minutes shows false regularly, and a slot abandoned six months ago shows false permanently. The column that distinguishes them is the last one: a slot whose distance from the current position keeps increasing is the one to worry about, whatever its connection state says. Note that the distance is not zero even here, seconds after creation, because the position a slot reserves is where a replica would have to begin replaying rather than where the log currently ends.

What to check on a real cluster

The distance above, tracked over time and per slot, is the alert worth having, together with a bound so an unattended slot cannot fill a disk. Which bound to set, what to do with a slot whose consumer is genuinely gone, and how logical slots differ once the cleanup horizon is involved are in replication lag and slot health. When the server acts on the bound itself, replication slot invalidation is what it records.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.