PostgreSQL 18
pg_aios answers only while you watch
For anyone who read that 18 added an I/O view and found nothing in it.
Reference page, revised in place. Last updated .
A view with no memory
Almost every monitoring surface PostgreSQL offers is a counter. You read it, you read it again later, and the difference is the thing you wanted. Counters survive being looked at rarely, they survive a scrape that fails, and they are what the whole apparatus of exporters and rate functions is built around.
The view PostgreSQL 18 added for asynchronous I/O is not one of those. It reports the requests that are outstanding at the instant you ask, and it forgets them the moment they finish. There is no total, no stats_reset, nothing accumulates. It is closer in spirit to the activity view, which shows you what is running now, than to anything in the statistics family, and reading it the way you read a counter produces a confident and wrong conclusion within about a second.
On 17 the question does not arise, because there is nothing to ask.
SELECT count(*) FROM pg_aios;
ERROR: relation "pg_aios" does not exist
LINE 1: SELECT count(*) FROM pg_aios;
^
The empty answer everybody gets first
Here is what almost everyone sees the first time, on a cluster that is not doing anything in particular.
SELECT count(*) AS handles_in_flight FROM pg_aios;
handles_in_flight
-------------------
1
(1 row)
One. Not zero, which would have been a cleaner story and a less honest one: something on the server had a request outstanding at the moment the question was asked, and a fraction of a second later it did not. Run the same count a dozen times on a quiet cluster and most answers are zero, some are one, and none of them mean anything. The number is a sample of a population you have not characterised.
This is the part that sends people away thinking the feature does not work. The view is doing exactly what it says. What is missing is a reason for anything to be in flight while you look.
Making the server busy while you ask
To see rows, something other than the session doing the asking has to be reading, because a session waiting on its own I/O is not in a position to run a query about it. The fixture below is a table a little larger than the buffer pool on this cluster, and a second connection opened from inside the same script so that the scan and the observation can happen at once.
CREATE EXTENSION dblink;
CREATE TABLE archive_pages (page_id bigint, payload text);
INSERT INTO archive_pages SELECT g, repeat('p', 400) FROM generate_series(1, 300000) AS g;
CHECKPOINT;
CREATE EXTENSION
CREATE TABLE
INSERT 0 300000
CHECKPOINT
The second connection is sent a sequential scan and is not waited on, so the observing session stays free to sample the view twice while the scan is still running.
SELECT dblink_connect('scanner', 'dbname=' || current_database());
SELECT dblink_send_query('scanner', 'SELECT count(*) FROM archive_pages WHERE payload LIKE ''%q%''');
SELECT pg_sleep(0.05);
SELECT state, operation, length, target, f_buffered, f_sync
FROM pg_aios ORDER BY io_id LIMIT 5;
SELECT pg_sleep(0.05);
SELECT state, operation, count(*) AS handles, sum(length) AS bytes
FROM pg_aios GROUP BY 1, 2 ORDER BY 1, 2;
SELECT * FROM dblink_get_result('scanner') AS matched(rows_found bigint);
SELECT dblink_disconnect('scanner');
dblink_connect
----------------
OK
(1 row)
dblink_send_query
-------------------
1
(1 row)
pg_sleep
----------
(1 row)
state | operation | length | target | f_buffered | f_sync
------------------+-----------+--------+--------+------------+--------
SUBMITTED | readv | 122880 | smgr | t | f
COMPLETED_SHARED | readv | 65536 | smgr | t | f
(2 rows)
pg_sleep
----------
(1 row)
state | operation | handles | bytes
-----------+-----------+---------+--------
SUBMITTED | readv | 3 | 368640
(1 row)
rows_found
------------
0
(1 row)
dblink_disconnect
-------------------
OK
(1 row)
Two samples fifty milliseconds apart, and they do not agree with each other. The first caught two handles, one submitted and one already finished but not yet reaped. The second caught three, all submitted, totalling three hundred and sixty thousand bytes of read requests outstanding at that instant. That disagreement is the view telling the truth about a thing that changes thousands of times a second.
The length column is the number worth reading first. One of those requests was a hundred and twenty thousand bytes, which is fifteen blocks fetched in a single operation. That is the asynchronous layer and the read combining working together, and it is visible here in a way it is visible nowhere else: the I/O statistics view will tell you afterwards how many bytes moved and in how many operations, but only this view shows you the shape of an individual request while it is in the air.
The flag columns are cheap to misread and worth one sentence each. A buffered request is going through the operating system’s page cache, which is the normal case; the alternative appears if the server is using a method that bypasses it. A sync flag distinguishes a durability operation from a data transfer. The target column says what kind of object the request is against, which matters when you are trying to work out whether a stall is on the log or on a relation.
Once the scan finished, the view emptied.
SELECT count(*) AS handles_after_the_scan FROM pg_aios;
handles_after_the_scan
------------------------
0
(1 row)
How to actually get a signal out of it
Everything above adds up to a practical rule, which is that this view is a profiling tool and not a metric source, and the mistake is to treat it as the second thing.
Sampling is the technique that fits it. Ask for the row count, or a breakdown by state, on a short interval during a window you already believe is interesting, and keep the samples. What you are measuring is not the I/O, which the statistics views already count properly, but the queue depth: how many requests the server is managing to keep outstanding at once. A workload that should be reading fast and is showing one or two handles at a time is not being limited by the storage, it is being limited by how much work the server is willing to issue before waiting, and that points at the concurrency ceiling rather than at the disks.
Two things make the sampling worth doing rather than skipping. It is the only view that shows request size directly, so it is where you confirm that combining is happening at all before you go looking for reasons a scan is slow. And the state breakdown separates requests the server has issued from requests that have come back and are waiting to be collected, which are different problems: the first is storage latency and the second is the backend being too busy to notice its own completions.
Two things argue for not bothering. On a cluster whose working set fits in memory there is almost never anything in flight, and a sampler will run for hours and find nothing. And the view is unindexed and unaggregated by design, so sampling it aggressively on a large server is itself work.
The honest summary is that this is the first PostgreSQL view that rewards being watched rather than scraped, and the monitoring stacks most fleets run are built entirely around scraping. Fitting it in means accepting that some questions are answered by a session somebody opens during an incident, which is a category most dashboards have quietly stopped making room for.
What the states are telling you
The state column carries most of the diagnostic value, and the two values in the capture above are the ones a sampler will see almost exclusively.
A submitted request has been handed to the storage layer and has not come back. A count of these is queue depth in the ordinary sense, and it is the number to compare against what the storage is capable of. A handful of submitted requests on a cluster that is supposedly saturating its disks is evidence that the bottleneck is upstream of the disks.
A request in the shared completed state has come back and is waiting to be collected by the backend that wanted it. A sample dominated by these says something different and more interesting: the storage is keeping up and the backends are not collecting fast enough, which happens when a backend is busy doing work between fetches rather than waiting on them. That is usually a healthy picture, and it is occasionally the signature of a server with more in flight than it has processes free to deal with.
The distinction matters because both look identical in every other view. Time spent waiting on a read shows up as a wait event of the same kind whether the read is slow or the collection is late, and the two want different responses: more storage throughput in the first case, more concurrency or fewer requests in the second.
Two practical notes for anyone building the sampler. Sample often and cheaply rather than rarely and thoroughly; a count grouped by state is enough for the signal above and costs almost nothing, while selecting every column of every row on a busy server is real work at an interval short enough to be useful. And keep the samples that returned nothing. A sampler that records only the non-empty results reports a permanently deep queue, because it has thrown away every moment when there was nothing in flight, which on most clusters is nearly all of them.
What it costs, and what it does not tell you
Reading the view is cheap and needs no extension, no restart and no setting. The second connection used above is an artefact of wanting one script to both cause and observe the load; in a real investigation the second session is a person with a terminal.
Two limits are worth stating plainly because neither is obvious. The view shows requests, not waits, so a backend blocked on something that is not I/O at all does not appear here and has to be found in the wait events. And nothing in it is attributable to a query: there is a process identifier on each row, which is enough to join to the activity view while the request is outstanding and useless a moment later. If the question is which statement is doing the reading, this is not the view that answers it, and the per-backend I/O counters are a better starting point.