Connection pooling
Postgres looks idle and the application is timing out
Most connection incidents are pooler incidents, and from inside the database they are nearly invisible. Whatever you use to watch Postgres has to watch the thing in front of it too.
The queue is not where you are looking
A Postgres backend is an operating-system process with its own memory, so a few thousand of them is not a configuration choice, it is a different machine. That is why a pooler exists, and it is also why the pooler becomes the part that runs out first. When it does, the database has nothing to report: the sessions it holds are few, healthy and mostly idle. Everything that is going wrong is happening to clients that have not reached it.
The confusion this produces is durable. Application dashboards show timeouts. Database dashboards show a server with capacity to spare. Both are accurate. The queue lives between them, in a process that frequently belongs to a different team and is often not in the monitoring inventory at all.
There is a second-order version of it too. Transaction-level pooling multiplexes clients onto shared backends, which is what makes it efficient and also what makes it lose anything a session was relying on being kept. Applications that were written against a session and then moved behind a transaction pool fail in ways that read as random until somebody works out what the mode changed.
Four things any pooler needs watched
The first is a discovery problem rather than a measurement problem, and it is the one that quietly defeats most setups.
It finds the pooler in the first place
Plenty of estates have no reliable list of what sits between the application and each database. A tool connected only to Postgres sees a calm server and no queue, because the queue is somewhere else.
It times the wait before the handover
The number that matters is how long a client sat holding nothing at all before it was given a server connection. That measurement exists inside the pooler and nowhere else in the stack.
It uses the ceiling that applies
Saturation means nothing against the database connection limit when the limit that binds is the pool size. Those two numbers are usually far apart and only one of them is reachable.
It knows what a pooling mode removes
Transaction-level pooling quietly withdraws things sessions rely on. Advice to switch modes that does not name what breaks is advice to take an outage on a Thursday.
Finding out what is actually in front
Before anything is measured, the question is answered by probing rather than by asking somebody. What responds on the port, what the pooler’s own administrative interface says when one exists, and what the sessions arriving at the database look like are enough to identify which pooler this is and in which mode it is running. The pooling dimension lists what gets recognised, including the managed proxies where you never get to see a process at all, and the honest answer of none.
With that established the measurements have somewhere to live: how deep the pool is, how long clients are waiting before handover, how many server connections are actually in use, and how close to the ceiling that puts you — the pool’s ceiling, which is the one that binds. Those are collected alongside the database signals rather than in a separate tool, so a saturation event and the statements running at the time are on the same timeline. What that timeline is built from is on the observability section.
Most of what saturates a pool is not the pool. Backends sitting idle inside open transactions hold their server connection while doing nothing, and a blocking chain does the same thing further down. Pool exhaustion is therefore often the last visible symptom of something that started as a lock, which is why the sessions behind it are kept rather than summarised, and why the story usually continues on the query side.
Reloading a pooler is one of the remediation templates, and it is a good example of what those carry: a declared safety class, a statement of whether it can be undone, a dry run before it executes and a probe afterwards that proves the change took effect rather than assuming it. It reaches a host only through the fail-closed gate, at whatever rung of the autonomy ladder that tenant has been granted, and the decision is signed into the ledger whether it was allowed or refused.
What a saturation event reads like
The pooler, the clients, the database and the cause on one screen, in that order, because that is the order the question gets asked in.
in front PgBouncer on both application hosts · transaction pooling clients waiting for a server connection, and have been for minutes pool at its configured size · the database is nowhere near its own server backends mostly idle in transaction, not running statements root one long idle-in-transaction session per host, holding its slot template reloading the pooler is reversible · a restart is not
Line three is the one that ends the argument between the two teams, and line five is the thing to fix. Reloading the pooler would clear the queue and those sessions would refill it within the hour. Note also that a pooler keeps pointing wherever it was told to point: after a failover, the cluster can be entirely healthy while the pool still holds connections to a node that is no longer the leader, which is a case for reading replication and pooling together.
The admin console will tell you most of this
Connection pooling and PgBouncer covers why a backend is expensive, what transaction pooling takes away, and how to read the admin console while clients are queueing. Lock contention and blocking trees is the other half, because the thing holding your pool open is usually holding a lock as well.
One application in front of one pooler is a problem you can hold in your head. It stops being that when the poolers differ per environment, when two of them are managed proxies you cannot log into, and when nobody can say from memory which databases sit behind which. At that point the inventory is the product, and fleet management is where that argument continues.
Find every pooler you own, including the ones nobody listed.
A pilot probes what sits in front of each database and measures the wait on the client side of it. Read-only from the first day, with nothing permitted to reload anything until you say so.