Question: Vacuum, bloat and wraparound
How do I know if autovacuum is keeping up?
Answered in the first paragraph. Last updated .
Compare how long ago each busy table was last vacuumed against how fast it earns new dead rows, and check whether the vacuums that did run were permitted to remove anything. The count of dead rows on its own answers nothing, because it climbs and falls on purpose. Autovacuum is behind when tables it should have reached are still waiting, or when it reaches them and reclaims nothing.
The same symptom, three unrelated causes
A table with a large and rising dead count can be in one of three states, and they need opposite responses.
It may never have qualified. Autovacuum triggers on a fraction of the table plus a fixed floor, so a very large table has to accumulate an enormous number of dead rows before it becomes eligible at all, and the per-table autovacuum_vacuum_scale_factor is the setting that decides this.
It may qualify constantly and never finish. One worker pinned to a huge table for hours leaves the rest of the schema unserved, and the cost delay that paces vacuum is applied per worker. Here the fix is throughput and worker count, and nothing about thresholds will help.
Or it may be running exactly as designed and forbidden from removing a single row, because something older is still entitled to see them. That is the xmin horizon, and it is the case where tuning autovacuum harder makes the server busier and the table no smaller.
What to look at instead
Age, not volume. The last autovacuum timestamp per table in pg_stat_user_tables, read next to that table’s write rate, tells you whether the schedule is being met. A table written to every minute and last vacuumed yesterday is behind regardless of what its estimated dead count says.
Concurrency next. Count the autovacuum workers in pg_stat_activity; if the maximum is in use most of the time, the server is saturated rather than idle, and raising thresholds will not create capacity.
Then the log. With log_autovacuum_min_duration set to something other than off, each run prints what it removed and, crucially, how many dead rows it could not remove yet and the oldest transaction that stopped it. That single line separates cause three from causes one and two without any other instrumentation.
The thresholds worth setting per table, and what to do once you know which of the three you have, are in autovacuum and table bloat. If the age of the table’s frozen id is what is rising rather than its dead count, the relevant behaviour is the vacuum failsafe, which abandons pacing entirely once the age gets dangerous.
For the version you are actually on, the autovacuum trigger calculator answers the question this page is about directly: whether the table is waiting on its threshold or losing on throughput, and how many index passes one vacuum takes at your maintenance_work_mem.