Question: Replication and slots
Why did my logical replication subscription stop?
Answered in the first paragraph. Last updated .
Almost always because the apply worker met a change it could not apply and raised an error. The subscriber’s log names the conflict, the relation and the transaction it was processing. Until somebody resolves it the subscription makes no progress, so its slot on the publisher stops advancing, and the publisher then keeps every log segment since that point. A stalled subscriber becomes a disk alert on the other machine.
The conflicts that stop it
A row that already exists on the subscriber is the most common, usually because something wrote locally to a table that is supposed to be replica-only. A missing row for an incoming update or delete is the mirror image. Both happen when a table is written on both sides, which is the arrangement logical replication makes easy and does not defend.
Schema is the other family. A column added on the publisher and not on the subscriber, a table added to the publication without the subscription being refreshed, or a type that does not match, all stop the worker in the same way. These are ordinary deployment ordering problems and they are why schema changes on a replicated pair have a required order.
Before PostgreSQL 15 a worker failing repeatedly left no counter anywhere and looked, from any dashboard, like a worker doing nothing much: the view that shows a subscriber looping is what made the state visible. PostgreSQL 18 goes further and counts each kind of conflict separately, including the ones that resolve themselves and previously left no trace at all.
Getting it moving again, best option first
Fix the data on the subscriber so the change applies. Delete the conflicting row, or create the missing one, and the worker succeeds on its next attempt. This is the only option that loses nothing.
If the transaction genuinely should not be applied, it can be skipped by naming its finishing position, which the error message gives you. Understand what that means: those changes are discarded and the two sides now differ in a way nothing will reconcile. Advancing the replication origin by hand is the blunter form of the same thing.
Then set the subscription to disable itself on error, so the next occurrence stops cleanly and visibly instead of a worker restarting into the same failure indefinitely.
Whatever you do, act on the publisher’s disk before the fix is agreed. The slot holding its position is the immediate risk, and dropping a replication slot safely covers what dropping it costs, which here is a full resynchronisation. Replication lag and slot health covers alerting on both sides at once.