Glossary: Connections
Connection storm
Also called: connection stampede, thundering herd.
Definition, revised in place. Last updated .
A connection storm is a feedback loop in which an application responds to database slowness by opening more connections, and the additional connections make the database slower still. Each PostgreSQL connection is an operating system process with its own memory and its own entry in structures the server scans, so past a certain point the cost of having connections exceeds the work they do. The loop ends when max_connections is reached and new logins are refused, which is usually the first thing anyone notices.
Why adding connections subtracts throughput
A database has a fixed amount of parallelism available: so many cores, so much disk concurrency, one lock on any given row. Connections beyond that do not create capacity, they create queueing, and PostgreSQL queues expensively compared with a pool that simply makes callers wait. Every backend contributes to the work of taking a snapshot, to the shared structures that track locks and transactions, and to memory pressure through its own work areas.
The loop needs an application that treats a timeout as a reason to retry on a new connection. A slow query holds a connection; the pool grows to compensate; the larger pool lengthens the queue; more requests time out; the pool grows again. The database is now spending its time on connection management, and every graph the application team looks at agrees that the database is slow, which it now is.
Two details decide how the incident presents. The reserved slots held back by superuser_reserved_connections are what let an operator in when everything else is refused, and a monitoring user that is not a superuser does not get them. And the failure is frequently not max_connections at all but the pooler’s own limit, or the operating system’s file descriptor limit, each with its own message and its own place to look.
Reading the headroom before it matters
The three numbers that bound the incident are one query, and the answer contains a small lie.
SELECT current_setting('max_connections')::int AS max_connections,
current_setting('superuser_reserved_connections')::int AS superuser_reserved,
count(*) AS client_backends
FROM pg_stat_activity
WHERE backend_type = 'client backend';
max_connections | superuser_reserved | client_backends
-----------------+--------------------+-----------------
100 | 3 | 1
(1 row)
The one client backend is the session asking the question, so the usable headroom on this server is the limit minus the reserved slots minus your own connection. That self-inclusion is harmless at one, and it is worth remembering when the same query is run by a monitoring agent that keeps its own connection open on every node.
The backend_type filter is what separates clients from the server’s own processes, which also appear in this view and are not subject to the limit. Counting rows without it inflates the answer by a handful on every server and by more on PostgreSQL 18, which runs additional I/O workers.
Breaking the loop
The durable fix is to stop the number of server connections from tracking the number of application threads, which is what a pooler does and what connection pooling and PgBouncer covers, including which pooling mode is compatible with the way your application uses transactions. If the connections are accumulating because each one is holding a transaction open rather than working, the state to look for is idle in transaction and the pool size is not the cause.