Question: Replication and slots
Why is my replication slot growing?
Answered in the first paragraph. Last updated .
Because the consumer behind it has stopped confirming how far it has read, and the primary is doing exactly what a slot asks: keeping every write-ahead log segment from that point onward so the consumer can resume. The slot is a promise, and nothing expires it by default. A disconnected replica, a paused change-data pipeline and a crashed subscriber all look identical from the primary.
Whether it is stalled or merely slow
These need different responses and the same column separates them. A slot that is not active has no consumer attached at all, and the amount retained will grow until something intervenes. A slot that is active but whose position advances more slowly than the primary writes is a throughput problem in the consumer, and it will catch up if the workload lets it.
Logical slots have a third case that physical ones do not. A logical slot cannot confirm past a transaction that has not committed yet, so one long-running write transaction on the publisher pins the slot no matter how fast the subscriber is. The subscriber looks healthy, the slot looks stuck, and nothing about the replica is wrong.
There is a cost beyond disk. A slot can also hold back the cleanup cutoff for the whole cluster, which is why an abandoned slot and a table that refuses to shrink are so often the same incident. That mechanism is the xmin horizon.
What stops it before the disk does
max_slot_wal_keep_size puts a ceiling on how much log any one slot may pin. When a slot crosses it the server sacrifices the slot rather than the cluster, and from PostgreSQL 18 a slot that has simply sat idle for too long can be invalidated on time rather than on volume. Both are deliberate, both destroy the slot, and both are better than a full disk on the primary.
Since PostgreSQL 14 the replication slot statistics view reports how much work each logical slot has spilled to disk, which is the signal that a slot is expensive rather than merely behind. From PostgreSQL 17 the slot view also records why a slot was invalidated, and how long it has been inactive, which is the first version where the question “when did this stop” has an answer on the primary.
Alert on an inactive slot with a growing retention rather than on the end state, because by the time the state is terminal the remedy is a fresh copy of the data. Replication lag and slot health covers the thresholds that fire early enough to act on.