Skip to content
dbexplore

Fleet management

The inventory was accurate the week it was written

Managing many Postgres clusters is mostly an argument about what you actually run. Every other question — which are at risk, which are drifting, which are about to need something — depends on that answer being current.

Everything downstream depends on the list

Ask any team running a lot of Postgres how many clusters they have and you will get a number, then a pause, then a qualification. The number came from somewhere — a tagging convention, an infrastructure repository, a page somebody keeps — and every one of those sources describes an intention rather than a state. The gap between the two is where the incidents are, because a cluster nobody has on the list is a cluster nobody has an alert on.

Scale changes the character of the work rather than its volume. One database is a thing you know. Twenty is a thing you can hold with discipline. Two hundred across four different Postgres platforms is a thing that can only be held by something that goes and looks, because the rate at which the estate changes exceeds the rate at which anybody updates a document about it.

And the platforms do not agree with each other. What counts as a node, what a failover does, whether you can install an extension, and which internals are yours to see all differ by platform. A fleet tool that flattens that away has made the estate look uniform, which is the one thing it definitely is not.

Four tests for anything managing many clusters

The fourth is the one nobody advertises, and it is the difference between a fleet view you trust at three in the morning and one you argue with.

The inventory is found, not typed

A list somebody maintained by hand was correct on the afternoon it was written. Anything that asks you to declare what you run has handed the hardest part of the job back to you.

One word, one meaning, every engine

Lag on a managed reader endpoint and lag on a pair you built are not the same measurement. Stacking them on a chart because they share a label produces a fleet view that is confidently wrong.

It looks again on a schedule

Estates change without a ticket. A node gets promoted, an extension appears, a proxy is introduced in front of a database. A description taken once at onboarding is an assumption by the second month.

It reports what it could not see

An absent signal is not a zero. Where an extension is missing or a managed platform does not expose something, the honest output is a gap with a name on it, not a healthy-looking blank.

How an estate gets described here

A cluster is characterised along seven dimensions before it is monitored: which Postgres it genuinely is, what decides leadership, what sits in front of it, how it is backed up, how it replicates, which extensions are present, and what is already watching it. That work is topology discovery, it is done by probing rather than by asking, and it is redone on a schedule so that a promotion, a new proxy or a removed extension arrives as an event with a date on it instead of as a discrepancy somebody trips over later.

Signals are then normalised against that description rather than against a common denominator. The same chart can carry two platforms because the discovery step established what each of them means by the number, and detection rules carry the engines they apply to, so a rule about internals that do not exist on a given platform never fires there. Where a measurement is impossible, the finding says so and the advice continues at reduced depth — visibly reduced, named in the output. The console side of that is one view over all of it.

On top sits the comparison nobody does by hand: configuration and schema on each cluster against a golden baseline for its class, so the ones that have quietly diverged are a list rather than an archaeology project. The per-cluster questions this makes answerable at estate scale have pages of their own — which tables cannot be cleaned, which followers would come up wrong, and which indexes nothing reads — and all three are questions whose answers stop agreeing with each other precisely at the scale this page is about.

A console is a good way to look at one cluster and a poor way to manage two hundred, which is why nothing here is console-only. The REST API answers the questions the interface answers — what you run, what is behaving oddly, which plans moved, what the advisor is proposing, what the ledger recorded, what remediation is waiting — so whatever already automates the rest of your infrastructure can ask them on a schedule rather than a person asking them on a Monday. The same answers come back over the MCP interface, which is what lets an assistant looking at one cluster and a job sweeping all of them work from one description of the fleet instead of two that quietly disagree.

Nothing here changes a cluster on its own. Every mutating action passes a gate that refuses by default, at the rung of the ladder that tenant has actually reached, and every decision lands in a signed append-only record. Who may do what to which class of cluster is the governance half of fleet work, and it is covered on the security monitoring page and in the policy gate section.

What a re-discovery pass reports

Not an inventory dump. The difference between what the estate was last time and what it is now.

fleet · discovery delta illustrative

The fourth and fifth lines are a pair, and they are the ones worth arguing about. An extension that disappeared during an upgrade silently reduces what can be advised on that cluster. Reporting the reduction is less flattering than staying quiet about it and considerably more useful, because otherwise the cluster looks like the one with no recommendations rather than the one nothing can be measured on.

Do the discovery pass by hand once

It is worth doing at least once on a cluster you think you know. Postgres topology discovery works out what is on the other end of a connection string from inside it — writer or replica, which nodes are siblings, what is preloaded, what is archiving — and pg_upgrade and extensions is the audit to run before a major version, which is where most estates discover what they were actually running.

Under a dozen clusters, a scheduled script that writes the answer somewhere durable is a genuinely good solution and costs you a day. What breaks it is not the number of clusters so much as the number of different Postgres platforms among them, because each one needs its own special case and the special cases are where the script rots.

Get a list of what you run that nobody had to type.

A pilot discovers the shape of every cluster you point it at and re-checks them on a schedule. Read-only throughout, so the first thing you get is an accurate description and nothing else.