Skip to content
dbexplore

PostgreSQL 15

The idle server that started writing logs

For anyone who has tried to tune checkpoints from counters alone because the log had nothing in it.

Reference page, revised in place. Last updated .

A default that is an opinion

Most default changes in a PostgreSQL release are adjustments to a number. Two of the defaults that changed in 15 are different in kind, because they change what the server writes to its log on a cluster where nobody has asked for anything.

Checkpoint logging went from off to on. Autovacuum logging went from off to logging any operation that takes longer than ten minutes. The Monitoring section of the PostgreSQL 15 release notes states both, and adds a warning that reads unusually bluntly for release notes: this will cause even an idle server to generate log output, which might cause problems on resource-constrained servers without log rotation.

That warning is the whole reason this is worth a page rather than a line in a table. It is simultaneously a good default and a genuine operational change, and the two facts have to be held at once.

Seeing the change

SELECT name, setting, unit, boot_val
FROM pg_settings
WHERE name IN ('log_checkpoints', 'log_autovacuum_min_duration')
ORDER BY name;

On 14:

            name             | setting | unit | boot_val 
-----------------------------+---------+------+----------
 log_autovacuum_min_duration | -1      | ms   | -1
 log_checkpoints             | off     |      | off
(2 rows)

On 15:

            name             | setting | unit | boot_val 
-----------------------------+---------+------+----------
 log_autovacuum_min_duration | 600000  | ms   | 600000
 log_checkpoints             | on      |      | on
(2 rows)

The autovacuum value is in milliseconds, so the new default is ten minutes, and the old value of minus one means the logging was disabled rather than set to zero.

Why these are the right defaults

Checkpoint logging is the single most useful thing in a PostgreSQL log for the class of problem that is hardest to diagnose from counters. The counters tell you how many checkpoints were timed and how many were requested, which distinguishes a cluster checkpointing on schedule from one being forced into checkpoints by write volume. What they do not tell you is how long each checkpoint took, how many buffers it wrote, how much it had to synchronise, and how far the write phase spread across the interval it was given.

Every one of those is in the log line, per checkpoint, with a timestamp. Tuning checkpoint spacing and the completion target from counters alone means inferring from aggregates what the log states directly. Anybody who has done both knows which is faster.

The autovacuum threshold is the same argument at a different scale. An autovacuum that finishes in seconds is noise. One that runs for more than ten minutes is either working on a large table, which you want to know about, or struggling, which you very much want to know about. Ten minutes is a reasonable line between those.

The reason these were off by default for so long is not that they were unhelpful. It is caution about log volume on small deployments, and 15 concluded that the diagnostic value wins.

The consequence the warning names

An idle cluster now writes a line every checkpoint interval, which by default means roughly every five minutes, forever. On one cluster that is nothing. On a fleet of several hundred small clusters shipping logs to a centralised store priced by ingestion, it is a line item that appeared without a change request.

Three things are worth checking after an upgrade, and they are cheap.

  • That log rotation is actually configured. A server writing to a file with no rotation and no retention now fills that file faster than it did.
  • That log shipping costs were sized for this. The volume increase is small per cluster and multiplies by cluster count.
  • That the checkpoint lines are reaching somewhere queryable. Doubling your log volume and then not using the new lines is the worst of both outcomes.

The third is the one that turns the cost into a benefit. These lines are worth parsing into a time series, because checkpoint duration and buffers written per checkpoint are exactly the signals the counters cannot give you, and the release that gave checkpoint counters a view of their own still does not give you per-checkpoint timing.

The upgrade trap, which is the opposite of the warning

Everything above assumes the new default is in force. On a large fraction of real upgrades it is not.

A cluster whose postgresql.conf was written years ago, has accumulated settings from several generations of advice, and is applied by a configuration management tool will very often contain an explicit setting of the value that was the old default. The server starts, the configuration wins, and the cluster runs PostgreSQL 15 with 14’s logging behaviour. Nobody notices, because nothing changed and nothing was supposed to.

The catalog distinguishes the two cases, and it is worth knowing how before an upgrade rather than after one.

ALTER SYSTEM SET log_checkpoints = off;
SELECT pg_reload_conf();
SELECT pg_sleep(1);
SELECT name, setting, boot_val, source FROM pg_settings WHERE name = 'log_checkpoints';
ALTER SYSTEM RESET log_checkpoints;
SELECT pg_reload_conf();
SELECT pg_sleep(1);
SELECT name, setting, boot_val, source FROM pg_settings WHERE name = 'log_checkpoints';
ALTER SYSTEM
 pg_reload_conf 
----------------
 t
(1 row)

 pg_sleep 
----------
 
(1 row)

      name       | setting | boot_val |       source       
-----------------+---------+----------+--------------------
 log_checkpoints | off     | on       | configuration file
(1 row)

ALTER SYSTEM
 pg_reload_conf 
----------------
 t
(1 row)

 pg_sleep 
----------
 
(1 row)

      name       | setting | boot_val | source  
-----------------+---------+----------+---------
 log_checkpoints | on      | on       | default
(1 row)

Three columns and they say three different things. The current value, the value the binaries ship with, and where the current value came from. A row where the setting differs from the shipped default and the source is a configuration file is a deliberate override, which may be exactly what you want. A fleet audit that only reads the current value cannot tell a cluster that has opted out from a cluster that never got the new default.

This generalises well beyond these two settings. Every default change in every release lands this way, and the question “is this cluster running the new default” is answered by comparing the current value against the shipped one, not by knowing which version is installed.

What is actually in the lines

It is worth knowing what the new output contains, because the decision about whether the volume is worth it depends entirely on whether the content is useful, and the content is better than its reputation.

A checkpoint produces two lines: one when it starts, naming what triggered it, and one when it completes. The completion line is the valuable one. It carries how many buffers were written and what share of the buffer pool that was, how many segments were added, removed and recycled, and the time spent in each phase of the checkpoint separately. The write phase and the synchronise phase are reported apart from each other, which matters because they fail differently: a long write phase is usually the completion target working as intended, and a long synchronise phase is storage that cannot keep up.

None of that is available from the counters. The counters say how many checkpoints happened and how they were triggered, and that is genuinely useful for deciding whether your segment budget is too small. They say nothing about duration, nothing about volume per checkpoint, and nothing about which phase the time went to.

The autovacuum line is the same kind of upgrade at a different scale. It reports pages scanned and skipped, tuples removed and remaining, the index passes performed, the buffer and write-ahead log activity, and the elapsed time. An autovacuum that took forty minutes and did three index passes is telling you its memory allowance was too small, and that conclusion is not reachable from any view.

The practical form is to parse both into a time series rather than to read them. Checkpoint duration, buffers per checkpoint and synchronise time make three panels that answer more questions about write behaviour than anything in the statistics views, and they arrive on a schedule whether or not anyone is looking.

What to do about the pair of them

For most fleets, nothing, which is the point of a good default. For fleets where log volume is a real constraint, the adjustment is not to turn checkpoint logging back off; it is to raise the autovacuum threshold, because that is the setting whose output scales with how busy the cluster is, while checkpoint lines arrive at a steady and predictable rate.

For fleets that were already setting both explicitly, the work is to decide again rather than to inherit. A setting that was chosen when the default was the other way around is a decision made in a different context, and an upgrade is the cheapest moment to revisit it.

What it costs

The logging costs log volume and a negligible amount of work at checkpoint time. The checkpoint line is written once per checkpoint by the process that just finished one, and the autovacuum line once per qualifying operation.

The real cost is downstream, in whatever ingests, stores and indexes the log, and it is proportional to cluster count rather than to cluster size. That is an unusual shape for a database cost and it is why the effect is most often noticed by a fleet of small clusters rather than by one large one.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.