Question: Replication and slots
How do I check replication lag in seconds?
Answered in the first paragraph. Last updated .
Two sources, and you want both. On the standby, the timestamp of the last replayed transaction compared with the current time gives a figure in seconds. On the primary, the replication view reports the write, flush and replay delays for each connected standby directly as intervals, already measured. The primary’s numbers are the better ones, and they have a blind spot the standby’s do not.
The trap in the standby’s own figure
It measures the age of the last transaction it replayed, not how far behind it is. On a primary that has gone quiet, nothing new arrives, the last replayed timestamp stops moving, and the calculated lag climbs by one second per second forever. Every team that alerts on this number alone gets paged the first weekend the write traffic stops.
The guard is to compare positions rather than clocks before trusting the seconds. If the last position received equals the last position replayed, the standby is current no matter what the timestamp arithmetic says, and the seconds figure should be reported as zero. Do that test first and the alert stops lying.
The blind spot on the primary, and what to alert on
The replication view only has rows for standbys that are currently connected. A standby that crashed, or whose network dropped, does not appear with enormous lag; it does not appear at all. So a rule written purely as a threshold on that view goes quiet at exactly the moment it should be loudest. Alert on the row being absent as well as on the value being high.
The three delays are not interchangeable and the distinction decides what a failure costs you: received, made durable, and actually applied are three different promises, which is write, flush and replay lag. A standby that is receiving and flushing but not replaying is safe for durability and useless for reads.
Watch bytes alongside seconds. Seconds is what a recovery objective is written in; the byte distance is what predicts the disk filling, and the two diverge whenever the standby is applying slowly rather than receiving slowly. When the distance is growing because a slot is holding segments rather than because a standby is behind, the page to read is why a replication slot grows. Replication lag and slot health covers turning all of this into alerting that survives a quiet Sunday.