Question: Replication and slots
ERROR: canceling statement due to conflict with recovery
Answered in the first paragraph. Last updated .
The query was running on a standby, and applying the primary’s changes would have destroyed something the query still needed. A standby is not allowed to stop replaying indefinitely, so after a grace period it cancels the query instead. Nothing is damaged and nothing needs repairing. You are being asked, in the form of an error, to choose between long reads and a replica that stays current.
What actually conflicts
The common case is cleanup. The primary vacuums away row versions that are genuinely dead there, the standby replays that, and a query on the standby is still entitled to see them under its own snapshot. There is no way to satisfy both, so one of them loses.
The others are less frequent and more abrupt. A strong lock taken on the primary by a schema change conflicts with anything reading that table on the standby. Dropping a tablespace conflicts with queries using it for temporary work. Dropping a database disconnects sessions attached to it. And there is a page-level variant that fires even when the rows being removed were not ones the query would have returned.
The detail that changes how you tune this: the grace period is measured from when the log data arrived, not from when your query started. A standby that is already behind has spent its allowance before your query even begins, so the same report that ran yesterday is cancelled in its first second today. That is why the symptom appears to be random and is actually a function of replay backlog.
The two settings, and the fact that you cannot have everything
The delay setting bounds how long replay will wait for queries, separately for streamed and archived data. Set to wait forever, the standby will serve any query and can fall arbitrarily far behind, which is correct for a reporting replica nobody would fail over to.
The feedback setting takes the other route: the standby tells the primary the oldest snapshot it is serving, and the primary’s vacuum holds back rather than removing those rows. Conflicts stop. The cost lands on the primary, where dead rows now accumulate for as long as the longest standby query runs, which is the xmin horizon extended across a network.
So there are three things on offer, current replay, long queries, and a primary that stays clean, and you may pick two. Deciding which replica is for reads and which is for failover, and configuring them differently, is usually better than one compromise on both. Replication lag and slot health covers the monitoring that tells you which one you actually built, and checking replication lag in seconds covers measuring the backlog that decides your grace period.