Skip to content
dbexplore

PostgreSQL 18

The I/O workers PostgreSQL 18 starts

For anyone whose process-count alert or backend_type panel is about to gain rows nobody added.

Reference page, revised in place. Last updated .

Three processes you did not configure

Start PostgreSQL 18 with a default configuration and count the processes. There are more of them than there were on 17, and nothing in your configuration file asked for any of them.

The reason is that 18 stopped issuing reads the way every previous major did. Until now, a backend that needed a block it did not have in the buffer pool called into the operating system and waited there until the data came back. That is simple, it is easy to reason about, and it wastes the interval between asking for a block and getting it, which on network-attached storage is most of the time the query spends. Version 18 introduces a layer that can issue several reads before waiting on any of them, and the default way it does that is to hand the requests to a small pool of dedicated processes.

Those processes are the visible part. The settings that control them are the part you have to know exists, because one of them cannot be changed without a restart and the release notes record it in a sentence about performance rather than a sentence about operations.

Asking both servers what they can be told to do

The settings arrived as a group, and the simplest way to see the group is to ask the server for it by prefix. On 17 the answer is one row, which is itself worth noticing.

SELECT name, setting, context FROM pg_settings WHERE name LIKE 'io\_%' ORDER BY name;
       name       | setting | context 
------------------+---------+---------
 io_combine_limit | 16      | user
(1 row)

The same question on 18 returns a family.

         name         | setting |  context   
----------------------+---------+------------
 io_combine_limit     | 16      | user
 io_max_combine_limit | 16      | postmaster
 io_max_concurrency   | 23      | postmaster
 io_method            | worker  | postmaster
 io_workers           | 3       | sighup
(5 rows)

Read the context column before the values. Three of the five cannot be changed without restarting the server, and one of those three is the choice of method itself. That makes the decision about how a cluster issues I/O a maintenance-window decision rather than a tuning decision, which is unusual for something introduced this recently and is the single most useful fact on this page.

io_max_concurrency is the one nobody mentions. It is not in the release note, and it is not really a default either: its boot value is minus one, which means the server computes the real figure at startup from the rest of the configuration. That makes it the only setting in the group whose value you cannot predict from a configuration file. On the cluster above, which runs a deliberately small buffer pool, it resolved to twenty-three. The same image started with the stock buffer pool resolved it to sixty-four. Two clusters you believe are identically configured can therefore be running with different ceilings on outstanding I/O, for a reason that appears in neither of their configuration files, and it is the first value to compare when two machines on the same storage behave differently.

io_method accepts three values. The default is the worker pool, and the other two are worth knowing by name so that a support conversation makes sense: one issues I/O synchronously in the requesting backend, which is what every earlier major did and is the setting to reach for if you suspect the new layer of something, and one uses the kernel’s own asynchronous interface where the platform has it.

The process list grows by three

Here is the same census of running processes, taken on each server with nothing else happening.

SELECT backend_type, count(*) AS processes
FROM pg_stat_activity
WHERE backend_type IS NOT NULL
GROUP BY 1 ORDER BY 1;
         backend_type         | processes 
------------------------------+-----------
 autovacuum launcher          |         1
 background writer            |         1
 checkpointer                 |         1
 client backend               |         1
 logical replication launcher |         1
 walwriter                    |         1
(6 rows)

And on 18:

         backend_type         | processes 
------------------------------+-----------
 autovacuum launcher          |         1
 background writer            |         1
 checkpointer                 |         1
 client backend               |         1
 io worker                    |         3
 logical replication launcher |         1
 walwriter                    |         1
(7 rows)

Three extra processes on an idle cluster, named io worker, present because io_workers defaults to three. Nothing about that is alarming in itself. What matters is the set of places where a process count is treated as a number with a meaning attached.

A connection-count alert that counts rows in the activity view without filtering by backend_type now sits three closer to its threshold on every cluster, forever. A capacity model that multiplies a process count by an assumed per-process memory footprint is now counting three processes that do not have that footprint. A panel that breaks the activity view down by process type gains a category, and if the panel is built on a fixed list of series rather than a group-by, the new category is silently dropped rather than shown. None of those is a failure the server can warn you about, and all of them are cheap to find before the upgrade rather than after.

The worker count is settable with a reload rather than a restart, which is the one part of this feature that can be tuned on a live cluster. Raising it is the documented response to storage that can serve more requests in parallel than three processes can keep in flight.

Asking a 17 server which method it uses does not get an answer, or even a null.

SHOW io_method;
ERROR:  unrecognized configuration parameter "io_method"

That error is the useful shape of this change for anyone writing a collector that has to work across majors: the setting is absent rather than defaulted, so a probe can branch on whether it resolves instead of comparing version numbers.

The rows that stay at zero

Now the part that surprises people, and the reason this page exists rather than a line in a release summary. The new process type gets rows in the I/O statistics view, exactly as every other process type does.

SELECT object, context, reads, writes, extends
FROM pg_stat_io
WHERE backend_type = 'io worker'
ORDER BY object, context;
    object     |  context  | reads | writes | extends 
---------------+-----------+-------+--------+---------
 relation      | bulkread  |     0 |      0 |        
 relation      | bulkwrite |     0 |      0 |       0
 relation      | init      |     0 |      0 |       0
 relation      | normal    |     0 |      0 |       0
 relation      | vacuum    |     0 |      0 |       0
 temp relation | normal    |     0 |      0 |       0
 wal           | init      |       |      0 |        
 wal           | normal    |     0 |      0 |        
(8 rows)

Eight rows, every counter zero. That is not a cluster that has done no reading. It is the accounting working the way it should: the I/O is attributed to the process that asked for it, not to the process that made the system call on its behalf. If it were attributed to the worker, every read on the cluster would collapse into three rows and the view would lose the ability to tell a vacuum from a customer query, which is most of what it is for.

Proving it takes a scan large enough to miss the buffer pool. This cluster has a small pool on purpose, the table is a little over a hundred megabytes, and the counters are cleared immediately before the scan so the numbers belong to it.

CREATE TABLE ledger_lines (line_id bigint, note text);
INSERT INTO ledger_lines SELECT g, repeat('l', 400) FROM generate_series(1, 300000) AS g;
CHECKPOINT;
SELECT pg_stat_reset_shared('io');
SELECT count(*) AS matched FROM ledger_lines WHERE note LIKE '%q%';
SELECT backend_type, context, reads, pg_size_pretty(read_bytes) AS volume
FROM pg_stat_io
WHERE reads > 0 AND object = 'relation'
ORDER BY reads DESC;
CREATE TABLE
INSERT 0 300000
CHECKPOINT
 pg_stat_reset_shared 
----------------------
 
(1 row)

 matched 
---------
       0
(1 row)

   backend_type    | context  | reads | volume 
-------------------+----------+-------+--------
 background worker | bulkread |  1310 | 83 MB
 client backend    | bulkread |   714 | 48 MB
 background worker | normal   |    66 | 568 kB
 client backend    | normal   |    32 | 256 kB
(4 rows)

Every byte is attributed to the client backend that ran the query and to the parallel workers it recruited. The worker rows stayed at zero throughout. So the three new rows in your backend_type breakdown will be there, they will be permanently empty, and anybody who notices them and concludes that the asynchronous layer is not being used will be wrong.

There is a second reading of that output worth keeping. The parallel workers did more of the reading than the leader did, which is ordinary, and it is the reason a per-process I/O panel keyed on client backend alone understates a parallel workload on any version. That is not new in 18. It is simply easier to notice now that there is another process type in the list to explain away.

Proving the layer is doing anything

A question follows directly from the rows above and has no obvious answer: if the worker processes never appear in the I/O accounting, how do you tell whether the asynchronous layer is being used at all?

Not from the process list. The workers are started whether or not anything uses them, so their presence proves only that the method is set to the default. Not from the I/O counters either, because the attribution above puts everything on the requesting backend either way, so the numbers look the same whichever method is in force.

Two things do answer it. The first is the size of the reads, which is visible in the I/O view as bytes divided by operations and which the page on read combining measures directly: a scan issuing reads much larger than a block is going through the new path, because the synchronous path has no way to produce them. The second is the view of requests in flight, which is empty on a server using the synchronous method and is the only place the layer is visible as itself rather than by its effects.

That matters when something goes wrong and the new layer is a suspect. The way to eliminate it is to set the method back to the synchronous one and see whether the symptom persists, and the awkwardness is that doing so needs a restart on a cluster that is by hypothesis already misbehaving. A bisect that costs a restart per step is one you want to do at most once, which is an argument for collecting the read sizes routinely rather than reaching for them during an incident.

There is a third method available on platforms with the right kernel support, and it is worth knowing it exists rather than testing it in production. It does the same job with fewer processes by using the operating system’s own asynchronous interface. The trade is that it bypasses the process pool entirely, so the worker count stops meaning anything, and the failure modes belong to the kernel rather than to PostgreSQL.

What to establish before the upgrade

The work here is short and none of it is urgent, which is exactly why it gets skipped.

  • Find every alert and dashboard that counts processes or connections without filtering on process type, and decide whether three is close enough to the threshold to matter. On a cluster sized for a hundred connections it is not. On one sized for twenty it is.
  • Decide the method deliberately rather than inheriting it. The default is reasonable, and the point is that changing it later needs a restart, so the cluster you would rather not restart is the one to decide about first.
  • Record what io_max_concurrency resolved to on each cluster, because it is derived at startup, it differs between machines you believe are identical, and it is the first thing to compare when two clusters on the same storage behave differently.

What it costs to look at any of this

Nothing on this page needs an extension, a restart or a privileged role beyond reading the statistics views. The process census is a view over shared memory. The I/O breakdown is a view over counters the server maintains whether anyone reads them or not.

The cost that is real is the one nobody measures: three additional processes exist on every cluster on the version, doing nothing at all on a cluster whose working set fits in memory. On a machine running one large database that is invisible. On a machine running dozens of small clusters side by side it is not nothing, and io_workers is the setting that answers it. Turning the method off entirely and going back to synchronous reads is available, needs a restart, and should be a decision you can point at a measurement for, because the default exists to make read-heavy work faster and usually does.

Put every Postgres you run on autopilot.

We onboard teams in small batches. Tell us about your fleet and we will reach out when a seat opens. One email, no drip campaign.