Question: Monitoring
Is pg_stat_statements safe to enable in production?
Answered in the first paragraph. Last updated .
Yes. It is the closest thing PostgreSQL has to a standard, it runs on an enormous number of production clusters, and the overhead is small enough that the question is usually the wrong way round: the risk of not having it is higher. The real costs are that loading it requires a restart, it claims a fixed amount of shared memory up front, and one setting decides whether the data is trustworthy.
The setting that decides whether the numbers mean anything
The extension tracks a bounded number of distinct statements. When a workload produces more than that, the least-used entries are evicted to make room, and if that happens continuously the view becomes a rolling window over whichever queries were busy most recently. Totals computed across it are then not totals of anything.
That churn is visible: the extension counts how often it has had to evict, and a counter that climbs steadily means the limit is too small for the workload. Since PostgreSQL 14 that counter is exposed in its own view, so the health of the collection can be monitored rather than assumed. Raising the limit costs memory allocated at startup, which is why the change needs a restart and why it is worth getting right in one go.
The second thing to decide is whether to record statements executed inside functions and procedures, or only the ones the application sent. Recording everything gives you the inner work and doubles the count of some statements when you sum; recording only the outer layer hides where the time went in a function-heavy schema. PostgreSQL 14 made that distinction visible per row rather than making it a guess.
What it is good at, and the gap it leaves
It is excellent at totals: which statement shape consumed the most time across the week, how that changed, which one is called far more often than anyone expected. PostgreSQL 17 added the read and write columns that separate a statement that is slow from a statement that is merely reading a great deal.
It is poor at moments. Because the rows are cumulative, it cannot tell you what was happening during the ten minutes the site was down, or what those sessions were waiting on. That is a different kind of collection and it is the subject of active session history.
One more caveat about grouping: two statements that differ only in their literals are one row, and that normalisation occasionally merges things you wanted apart. Query fingerprints covers what is being folded together and when the fold misleads.