Learn
The Postgres failures worth recognising on sight
Thirteen reference pages, each on one problem, each with the queries we would actually run against a cluster we had never seen before.
Most Postgres incidents are not novel. A handful of failure shapes account for the overwhelming majority of the pages an on-call database engineer takes, and each of them announces itself in the catalog long before it announces itself to a customer. The gap between the two is where these pages live.
Every guide below follows the same route: what the thing actually is, the query that shows whether you have it, the change that ends it, and the version caveat where one exists. The SQL is written for a stock server and the public statistics views, so it runs on a managed instance with no superuser as readily as on a box you own. Where a column moved between major versions, the page says which version it is talking about, because a monitoring query that silently returns nothing is worse than one that errors.
These are not release notes. Read the blog for dated engineering notes with an argument in them; read these when you have a cluster in front of you and a question about it. If your question is about our software rather than about Postgres — what it connects as, what it takes, what it is allowed to change — that is security and trust, where what leaves your database is enumerated on both sides of the boundary.
-
Vacuum
Autovacuum and table bloat
For anyone watching a table's disk footprint grow while its row count stays flat.
-
Freezing
Transaction ID wraparound
For anyone who has seen the warning about vacuuming within some number of transactions and wants to know how worried to be.
-
Replication
Replication lag and slot health
For anyone whose standby is behind and who needs to know whether that is a network problem, a replay problem or a disk problem.
-
Pooling
Connection pooling and PgBouncer
For anyone whose application reports timeouts while the database itself looks almost idle.
-
Concurrency
Lock contention and blocking trees
For anyone staring at a wall of sessions waiting on a lock and trying to find the one that started it.
-
Planning
Query plan regression
For anyone being told the application is slower while every dashboard says the database is fine.
-
Indexing
Unused and missing indexes
For anyone about to drop an index on a production table and wanting to be sure first.
-
Wait events
Active session history in Postgres
For anyone asked what the database was doing at 03:10 last Tuesday and having no way to answer.
-
High availability
Config drift between primary and standby
For anyone whose failover plan assumes the replica is configured like the thing it is replacing.
-
Version 18
What PostgreSQL 18 changed in monitoring
For anyone about to upgrade who would rather find the broken panels before the first incident on the new version.
-
Upgrades
pg_upgrade and extensions
For anyone planning a major version upgrade on a cluster that has been alive long enough to accumulate things.
-
Discovery
Postgres topology discovery
For anyone handed a connection string and expected to say what is on the other end of it.
-
Agents
Postgres observability for AI agents
For anyone who has just given a model read access to a production database and would like to know what it is doing in there.
Running more than one of these? The same signals across a fleet are what the platform is built to watch, on every shape of cluster topology discovery can find, for whichever Postgres you happen to run. The indexing guide above has a companion on the fleet side of that question: the index advisor, and what it attaches to a recommendation before anyone is asked to approve one. So do three others: vacuum, query performance and replication, each picking up where the guide above it stops being something one person can do by hand.
Put every Postgres you run on autopilot.
We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.