AI agent
Letting an agent near production is a governance problem
The hard part was never getting a model to write plausible SQL against your fleet. It is deciding what happens on the occasion it is confidently wrong, and arranging that before it runs rather than afterwards.
Capability is not the constraint any more
A current model handed catalog access and a question about a slow statement will do respectable work. It will read the plan, notice the scan, propose something sensible, and explain itself in a paragraph your manager can follow. Anybody who has tried it knows the demonstration is genuinely impressive, and anybody responsible for the database also knows that the demonstration is not the part they are worried about.
The worry is the tail. A model that is right forty-nine times out of fifty and confidently wrong on the fiftieth is a fine research assistant and an unacceptable operator, because the fiftieth occasion is a production database and nobody was watching that one particularly closely. The instinct to solve this by writing a stern instruction into the prompt is understandable and does not work: the instruction arrives through the same channel as the request, and the failure you are guarding against is the model being wrong about which of the two it is reading.
So the interesting question is not how capable the agent is. It is what the agent is structurally unable to do, who decided that, where the decision is written down, and what is left behind afterwards for somebody who was not in the room.
Four things to require before you connect one
Ask these of any vendor putting an agent near your data, ourselves included. The second one is the one that gets answered vaguely.
The limits live outside the prompt
Telling a model what not to do is a preference expressed in the same channel as the request. A rule evaluated before the call is admitted is a limit. Only one of those survives a novel phrasing.
Silence counts as refusal
Deny, timeout, unreachable, malformed answer, no rule written for this case at all: if the thing that decides cannot say yes, the answer is no. An evaluator that is down must stop the estate changing, not wave it through.
Trust is earned per class of action
Not a setting somebody ticks on Friday afternoon. A class of action moves up only on its own record, and there have to be classes that never move up however good the record gets.
A sceptic can audit it afterwards
What was proposed, what was allowed, what was refused and what actually ran, in a record whose integrity you can check without taking the vendor’s word for any of it.
What ours may do, and what it may not
DBExplore exposes its capabilities over the Model Context Protocol, so Claude, Cursor or an assistant you wrote yourself can investigate a slow statement, read a range of the audit record or put forward a remediation. The tool families are declared, the connection is authenticated and scoped to one tenant, and the default is read-only. There is no undeclared tool and no privileged route that skips what a person would meet — the MCP section lists the surface.
Questions asked in plain language take the same path. The model drafts, a gate verifies the draft only reads before it is allowed to run, and you confirm anything that would do more than read. The division of labour is deliberate: the model is good at proposing and bad at being certain, so it proposes, and something that cannot be talked out of its position decides.
That something is a policy engine evaluated before the call is admitted, and it is fail-closed. Anything that is not an explicit allow — a denial, a timeout, an unreachable evaluator, a malformed response, a case nobody wrote a rule for — is a refusal. Every tenant begins where nothing mutates at all. A class of action moves up the ladder only on a measured record for that class in that tenant, the disruptive classes require more than one approver and are never eligible for automatic application, and there is a kill switch that stops the estate’s action plane outright. The gate section has the rest.
Everything either decision produces is appended to a signed, chain-hashed ledger per tenant — proposals, approvals, refusals and runs alike — which an auditor can verify without taking our word for it. The whole thing can also run inside your own boundary, up to and including an air-gapped deployment, which is the answer for teams whose objection to an agent is about where the data goes rather than about what the agent does. Security monitoring covers the connection model that sits under all of it.
A refusal, in full
The interesting output is not the one that succeeded. It is what comes back when an assistant asks for something it has not been granted.
tool remediation · cancel a long-running statement on one cluster caller an assistant acting for a named operator, tenant-scoped gate evaluated before the tool runs, against rules you wrote down result refused · this tenant holds the rung below one-click for this class returned the proposal, the evidence, and the rule that refused it ledger the refusal is signed into the record, exactly as an approval is
Line five matters as much as line four. A refusal that returns nothing teaches the operator nothing and invites somebody to go around it; a refusal that hands back the proposal, the evidence behind it and the rule that stopped it turns into a conversation about whether the rule is right. That is the outcome we want, because the rule is the thing you should be arguing with. The evidence itself is assembled the same way it is for a person, which is the subject of query performance.
Where we have stopped on purpose
We hold ourselves to the same ladder we are describing, and we wrote down why in why we capped ourselves at rung 2. The short version is that the precision record required to justify a higher rung has to be measured on real estates, and until it exists the honest position is that our agent recommends and a person decides. We would rather say that on a landing page than discover we had implied otherwise on a call.
If what you want is the agent reading across an estate rather than a single cluster, the description it reasons over is the one built by fleet discovery, and it is as accurate as that pass was. An agent cannot be better informed than the inventory underneath it, which is an unglamorous thing to put on a page about artificial intelligence and the first thing that actually limits one.
None of this is the only question worth asking about an agent and a database. The other one is what the database sees when an agent is the client rather than a deployed application, which changes the statement statistics, the connection pattern and the signal you would have relied on to notice a slow query. That is a Postgres problem rather than a product one, and it has its own reference page: observability for AI agents.
Point your own assistant at it and try to get past the gate.
A pilot starts with the read-only tools and a rung that mutates nothing. The interesting part of the evaluation is what gets refused, so bring the assistant you already use.