Aurora's writer endpoint can lie to you
CloudWatch measures instances. Your application talks to endpoints. For the minute those disagree, every instance metric looks healthy while writes fail.
There is a particular kind of incident that makes people distrust their monitoring for months afterwards. The application is throwing write errors. Someone opens the database dashboard. Every instance is green. CPU is fine, memory is fine, replica lag is fine, no instance is restarting. The dashboard and the application cannot both be right, and for a minute or two nobody can work out which one is lying.
On Aurora, both are telling the truth. They are answering different questions.
Instances and endpoints are not the same thing
Aurora separates compute from storage, and it gives you endpoints rather than hostnames. The writer endpoint is a DNS record that points at whichever instance currently holds the writer role. The reader endpoint round-robins across the instances currently in the reader rotation.
CloudWatch, meanwhile, publishes metrics per instance. CPU on this instance, connections on that one, replica lag on the third.
Most of the time this distinction is invisible, because the endpoint points where you would expect and the instance behind it is healthy. During a failover it stops being invisible.
When Aurora promotes a replica, the writer endpoint has to be repointed, and DNS has to propagate. For the window in between, the writer endpoint can still resolve to an instance that has already given up the writer role and is no longer accepting writes. Your application connects successfully, because the instance is up and listening. It then fails on the first write, because that instance is now a reader.
Every per-instance metric during that window is genuinely healthy. The old writer is a healthy reader. The new writer is a healthy writer. Nothing is down. The only thing that is broken is the mapping between them, and the mapping is not a metric.
Why lag skew alerts fire on the wrong nodes
The same confusion shows up in a quieter form.
A common alert is skew across replicas: if one reader is much further behind than the others, something is wrong with it. Reasonable rule, and on a plain Postgres cluster it works.
On Aurora it produces false positives, because a synchronous standby publishes replica lag like any other instance while not being in the reader endpoint rotation at all. No client reads from it. Its lag is real and also entirely irrelevant to anybody’s query latency, and an alert comparing it against instances that do serve reads is comparing things that are not comparable.
The fix is not a smarter threshold. It is scoping the comparison to the instances actually in the rotation, which means knowing the rotation, which means asking the endpoint rather than the instance list.
Units, while we are here
One more trap that costs people an afternoon. Aurora reports cluster replica lag in milliseconds through CloudWatch, but the per-instance lag you can read from inside Postgres is a time interval, and some tooling normalises one and not the other.
The result is a dashboard where the same lag appears twice, three orders of magnitude apart, and whichever number someone happens to look at determines whether they think the cluster is fine. We normalise everything to milliseconds on ingest for exactly this reason, and we would recommend picking a unit and enforcing it at the boundary rather than at the chart.
What to actually measure
If you take one thing from this, make it the first item.
- Probe the writer endpoint for write-ability, not just reachability. Open a connection and confirm the server will accept a write. Reachable and writable are different states, and the gap between them is exactly the failover window.
- Scope reader comparisons to the rotation, so that instances nobody reads from cannot trigger an alert about read latency.
- Track the volume and capacity ceilings, which are cluster properties rather than instance properties and therefore absent from a per-instance view. Serverless capacity against its maximum belongs here too, because hitting the ceiling looks like a slow database rather than a limit.
- Alert on the topology changing, not only on instances degrading. A failover that completes cleanly is still something you want to know happened, because it usually has a cause worth understanding.
The general shape of the mistake
Aurora is the clearest example, but the pattern is not specific to it.
Anywhere a managed platform puts an abstraction between your application and the process serving it, your monitoring has a choice: measure the thing the platform bills you for, or measure the thing your application talks to. Most tools measure the former, because that is what the cloud provider’s metrics API offers and it requires no probing.
The related discipline, making sure the node that gets promoted is configured like the one it replaces, is covered in config drift between primary and standby.
The trouble is that the abstraction is precisely where the interesting failures live. The instance is healthy, the storage is healthy, the cluster is healthy, and the path between them has a gap in it for ninety seconds. If nothing in your stack is measuring the path, nobody finds out until the application does.