Replication monitoring
The standby is behind, which is three different sentences
Behind on receive, behind on replay, and holding write-ahead log for something that stopped listening are separate failures with one symptom. Deciding which one you have is most of the work.
One symptom, several unrelated emergencies
A follower that cannot keep up with the network is a capacity problem, and it is the one everybody assumes. A follower that is receiving everything on time and applying it slowly is usually a long read holding replay back, which is a workload problem on the follower and has nothing to do with bandwidth. A slot with nothing attached to it is a third thing entirely: the leader is fine, every follower is current, and a volume is filling up on a clock nobody is watching.
These are distinguished by which catalog you read and from which node you read it, and that is precisely what is awkward to arrange. Most monitoring is pointed at a writer, because a writer is what the application connects to and what somebody put in a connection string. The measurements that separate the three live on the other machines.
Then there is the failure that only appears during the event you built replication for. A follower carrying a smaller setting than the node it is meant to replace will start, accept connections, and behave differently under load. Nobody finds that in normal operation, because in normal operation nothing asks the follower to be the leader.
Four questions to put to a replication monitor
Worth asking of the exporter you already run, and of us. The last one is where most setups turn out to be blind.
It asks the standby, not only the leader
A leader knows what it sent. Only the follower knows what it has replayed, and replay is the number a reader is actually waiting on. One of those is easy to collect and the other one is the truth.
It keeps the three numbers apart
Bytes behind, seconds behind and write-ahead log retained for a slot measure different failures. Paging on the wrong one wakes somebody for a quiet night and stays silent for a disk filling up.
It knows which kind this is
Physical streaming, native logical, a replication extension and a distributed engine report progress in different units and fail in different ways. One lag chart over all four is a chart about nothing.
It watches the slot itself
An abandoned slot holds write-ahead log on the primary until the volume is full, and the lag graph stays flat the entire way there because nothing is behind. Nothing is connected at all.
What we collect, and from which node
Every node in a cluster is read in its own right, and the readings are assembled into one picture of the cluster rather than a row per connection string. Receive and replay are kept as separate measurements on each follower; slot state and retained log are read where the slot lives; subscription conflicts are read where the subscription is. Cross-region paths that are invisible from the primary side appear because the secondary was asked directly. The replication dimension lists the paths that get recognised, from native streaming through logical extensions to distributed engines, and each is measured in the units that engine genuinely reports rather than converted into a house unit that means nothing.
Who is leading is a separate question from who is behind, and it is answered through whatever is making that decision — an orchestrator where one is running, the configuration store directly where it is not. The high-availability dimension sets out what is recognised. On top of it sit the signatures for the ways this goes wrong as a cluster rather than as a node: a leader that will not settle, failovers arriving in a burst, a synchronous follower stuck in a way that stalls commits on the leader, and a quorum that no longer exists.
Settings on each node are compared against the others continuously, so a follower that would come up weaker than the leader is a finding on an ordinary Tuesday instead of a surprise during the failover. The shape of the cluster is itself re-checked on a schedule, which means a node added, removed or promoted registers as an event — the same mechanism that keeps an estate of clusters described correctly rather than described once.
Fixing any of it is governed. Dropping an abandoned slot reclaims the volume and permanently ends that subscriber’s ability to catch up, which makes it disruptive by classification: more than one approver, never eligible for automatic application, and recorded either way in a signed append-only ledger you can verify yourself. The policy gate sets out how that refusal works when it is the gate itself that is unreachable.
A cluster summary, not a lag number
Six lines that let somebody on call decide what this is before opening anything.
cluster eu-west · one leader, two followers, one logical subscriber replay follower-b is behind on replay while receive is current cause a long read on follower-b is deferring replay, not the network slot the logical slot is inactive · retained log growing on the leader drift follower-b would come up with a lower worker setting than the leader action dropping a slot is disruptive · approval required, never automatic
Two independent problems are visible there and only one of them is urgent tonight. The deferred replay is a workload decision somebody can make in the morning. The inactive slot is a volume filling up. A failover in the middle of that is also the moment the pooler in front of the cluster matters more than the cluster does, which is its own subject.
The catalogs are open to you
Replication lag and slot health has the queries that split the three measurements apart, which to page on, and how a slot ends a cluster. Config drift between primary and standby lists the settings a follower must not carry lower than its leader and how to diff two nodes without fooling yourself. Run them against your own pair this week.
For one cluster with two followers that is enough, and we would rather you kept the money. It stops being enough at the point where the answer depends on which orchestrator is in charge of a particular cluster and nobody on the rota knows which that is without looking.
Find out what your followers would come up as.
A pilot reads every node in the cluster, not the writer. You see replay, slot retention and node-to-node drift on your own estate before anything is allowed to change a setting.