Question: Replication and slots
What happens if a synchronous standby goes down?
Answered in the first paragraph. Last updated .
Commits stop. The documentation puts it bluntly: transaction commits may never complete if a required synchronous standby crashes. The transaction is already written locally, and it is waiting for an acknowledgement that will not arrive. Reads carry on normally, so the symptom is a database that answers every query and accepts no writes, which looks like a hang and is the configuration doing precisely what it was told.
Why there is no timeout
There deliberately is not one. A timeout would mean the cluster silently downgraded its durability guarantee at the exact moment the guarantee was being tested, and nobody would know which transactions were covered and which were not. Refusing to proceed is the honest failure, and it is the reason this mode is a choice rather than a default.
The documented way out is to change the configuration rather than to wait. Reduce the number of standbys commits must wait for, or clear the list entirely, and reload on the primary; the waiting transactions then complete. It takes seconds and it is an explicit, logged decision to accept the reduced guarantee, which is the right shape for that decision.
A fast shutdown also releases waiters, which is worth knowing before somebody reaches for a harder kill.
Configuring it so a single failure is survivable
Name more standbys than you require. A quorum form that waits for any one of three is available whenever one of the three is up, where a priority form naming a single standby has no redundancy at all. This is the whole difference between a durability setting that improves availability and one that halves it.
Monitor the state of each standby in the replication view rather than only the lag, because the field that says whether a standby is currently counted as synchronous is what predicts this outage. A standby that has quietly fallen out of the set has already removed your margin.
Then decide whether you needed this mode at all. Waiting for a remote flush is a strong guarantee with a real availability cost; the asynchronous default loses a small window of transactions in a crash and never blocks. Synchronous commit covers the levels between those two, and can synchronous_commit off lose data covers the other end of the same dial. Replication lag and slot health covers the alerting that catches a standby leaving the set before a commit does.