Glossary: Vacuum and freezing
Bloat: the space that is not holding data
Also called: table bloat, index bloat, table fragmentation.
Definition, revised in place. Last updated .
Bloat is the difference between the size of a relation on disk and the size its currently visible rows would need. It comes from row versions that were superseded, removed by vacuum, and left as gaps that later rows may or may not fill. A table with some bloat is a table that has room to absorb the next round of updates without extending its file, which is why the target is never zero and why a percentage on its own decides nothing.
Three things wear the same label
The word covers at least three situations that want different responses, and a single number on a dashboard hides which one you have.
Free space inside a table that is being reused is not a problem. It is the working set of an update-heavy table doing exactly what it should, and shrinking it just means the next thousand updates extend the file again.
Free space that is never reused is a problem, and the usual cause is that the gaps are in the wrong places: a table whose rows are inserted in one order and updated in another accumulates holes too small for the rows that arrive later. That one shows up as a file that grows steadily while the live row count is flat.
Index bloat behaves differently again. An index does not reuse a freed entry unless a new key lands in the same page range, so an index on a monotonically increasing column can keep growing while the table it points at stays level. Indexes are also the part that a table rewrite fixes for free and a routine vacuum mostly does not, which is why the two numbers drift apart over months.
Where the number you are reading comes from
Almost every bloat figure in circulation is an estimate. The widely copied estimate queries multiply the row count in the catalog by an average row width derived from the planner’s column statistics, then compare the result against the page count. Both inputs are sampled rather than counted, so the answer inherits whatever error the last analyze introduced, and on a table with variable-width columns or a lot of nulls the estimate can be out by a wide margin in either direction.
That is a good enough instrument for ranking tables against one another, and a bad one for deciding to take an exclusive lock on the largest table in the cluster. Before a rewrite, measure. The difference between estimating and counting, and the extension that does the counting, is set out in autovacuum and table bloat.
What to do once you know which kind it is
Finding the cause comes before shrinking anything, because a rewrite without a cause is a rewrite you will repeat next quarter. The three reasons dead rows accumulate, the counters that separate them and the per-table settings that fix the common one are all in that same guide. If the count refuses to fall no matter how often vacuum runs, the thing to read first is the xmin horizon, which is the mechanism that forbids the cleanup rather than a sign the cleanup is behind.