Skip to content
dbexplore

Learn

The Postgres failures worth recognising on sight

Thirteen reference pages, each on one problem, each with the queries we would actually run against a cluster we had never seen before.

Most Postgres incidents are not novel. A handful of failure shapes account for the overwhelming majority of the pages an on-call database engineer takes, and each of them announces itself in the catalog long before it announces itself to a customer. The gap between the two is where these pages live.

Every guide below follows the same route: what the thing actually is, the query that shows whether you have it, the change that ends it, and the version caveat where one exists. The SQL is written for a stock server and the public statistics views, so it runs on a managed instance with no superuser as readily as on a box you own. Where a column moved between major versions, the page says which version it is talking about, because a monitoring query that silently returns nothing is worse than one that errors.

These are not release notes. Read the blog for dated engineering notes with an argument in them; read these when you have a cluster in front of you and a question about it. If your question is about our software rather than about Postgres — what it connects as, what it takes, what it is allowed to change — that is security and trust, where what leaves your database is enumerated on both sides of the boundary.

  1. Vacuum

    Autovacuum and table bloat

    For anyone watching a table's disk footprint grow while its row count stays flat.

  2. Freezing

    Transaction ID wraparound

    For anyone who has seen the warning about vacuuming within some number of transactions and wants to know how worried to be.

  3. Replication

    Replication lag and slot health

    For anyone whose standby is behind and who needs to know whether that is a network problem, a replay problem or a disk problem.

  4. Pooling

    Connection pooling and PgBouncer

    For anyone whose application reports timeouts while the database itself looks almost idle.

  5. Concurrency

    Lock contention and blocking trees

    For anyone staring at a wall of sessions waiting on a lock and trying to find the one that started it.

  6. Planning

    Query plan regression

    For anyone being told the application is slower while every dashboard says the database is fine.

  7. Indexing

    Unused and missing indexes

    For anyone about to drop an index on a production table and wanting to be sure first.

  8. Wait events

    Active session history in Postgres

    For anyone asked what the database was doing at 03:10 last Tuesday and having no way to answer.

  9. High availability

    Config drift between primary and standby

    For anyone whose failover plan assumes the replica is configured like the thing it is replacing.

  10. Version 18

    What PostgreSQL 18 changed in monitoring

    For anyone about to upgrade who would rather find the broken panels before the first incident on the new version.

  11. Upgrades

    pg_upgrade and extensions

    For anyone planning a major version upgrade on a cluster that has been alive long enough to accumulate things.

  12. Discovery

    Postgres topology discovery

    For anyone handed a connection string and expected to say what is on the other end of it.

  13. Agents

    Postgres observability for AI agents

    For anyone who has just given a model read access to a production database and would like to know what it is doing in there.

Running more than one of these? The same signals across a fleet are what the platform is built to watch, on every shape of cluster topology discovery can find, for whichever Postgres you happen to run. The indexing guide above has a companion on the fleet side of that question: the index advisor, and what it attaches to a recommendation before anyone is asked to approve one. So do three others: vacuum, query performance and replication, each picking up where the guide above it stops being something one person can do by hand.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.