Why we capped ourselves at rung 2
We built an action plane and a fail-closed policy gate, then refused to let any action run unattended until it could prove it was reversible. The ladder.
There is a moment in every database automation project where somebody asks the question you have been avoiding. Ours came in July, in a review of a design document that was supposed to be about naming. We had the action adapters, the policy rules, the anomaly library, an SLO engine and an advisor, and all of it sat behind a confirmation flag because nobody had written down how much of a customer’s database the product was allowed to touch on its own.
That is not a mechanism problem. The mechanism was done. It was a values problem, and values problems do not get fixed by shipping more code. So we wrote the ladder.
Four rungs, each a superset of the last
An action’s rung is a property of the action. Not of the customer, not of the incident, not of how confident the model feels that afternoon.
Rung 0 is observe. Detect and explain, offer nothing. Rung 1 is recommend: a specific action with its dry-run output and predicted effect, and a human decides and runs it. Rung 2 is one-click. The action arrives pre-validated with its rollback already computed, and a human approves it in one gesture. Rung 3 is auto-apply. It runs under policy, watches the result, and rolls back on regression while the human audits afterwards.
Each rung adds a precondition. Rung 1 needs a working dry-run. Rung 2 needs a declared, tested verify step and a pre-computed rollback. Rung 3 needs declared reversibility, an implemented rollback, and a policy that denies by default.
We shipped with the ceiling at rung 2. Here is why.
There are no read-only actions
The first thing we found when we classified the adapters was that every one of them mutates something. Some are local mutations, some touch the whole cluster, some are disruptive. There is no adapter you can justify at rung 3 by saying it cannot hurt anything. Every rung 3 argument has to be a reversibility argument.
That reframed the whole exercise. Safety class and reversibility are not the same axis. Cancelling a query and terminating an idle session are local, low-blast-radius actions that destroy in-flight user work and cannot un-destroy it. Adjusting a serverless autosuspend timeout is the same safety class and can be put back exactly as it was. Same tier. Opposite reversibility.
Reversible is a fact about the world. Rollback is a fact about our code.
The first draft treated rollback as implicit. If an action is reversible, surely we can reverse it. We rewrote that paragraph a day later, and the rewrite is the most important sentence in the document.
Reversible describes the world: the prior state can, in principle, be restored. Rollback describes our code: a promise that we will restore it, unattended, without a human remembering how. Those are different claims, and a governor with nothing to invoke on regression is worse than no autonomy at all, because it has told the operator to stop watching.
When we applied that standard, fewer adapters qualified for the top rung than we had assumed, and we chose to gate on it anyway. The durable fix was a capture contract. Before an adapter mutates anything, it snapshots the prior value, and an audit test asserts that every adapter with a live execute path also has a live rollback path. The requirement arrives with the capability instead of being remembered later by whoever is on call.
What the gate actually refuses
The ladder only means something because the gate underneath it fails closed. Before a mutating action runs, a policy engine evaluates it. Every outcome that is not an explicit allow, including deny, timeout, unreachable, a malformed response or a policy that does not exist yet, coerces to deny. If the gate is down, automation is down. Your database is not.
On top of the gate sits the approver rule. Tenants start at approve. Cluster-wide and disruptive actions require more than one distinct human and can never be switched to auto, and that check runs before the runtime even starts. Then verify-after-act probes the live system so a session that had already gone away is reported as a no-op rather than a success. Then the proof is signed into the ledger.
Destructive operations get no rung at all. There is no adapter for dropping a table, truncating one, or a planned failover, so they cannot be requested, by a human or a model.
Where the bar is still open
The open questions are more interesting than the closed ones. Who owns reversibility classification? Our answer is that reversibility defaults to false and each adapter opts in with a written justification, so the unsafe default is the silent one. Should rung 2 require a rollback that has been executed in test, not merely computed? We think yes, and we are moving the bar there.
And there is the health score, which rolls SLO attainment, anomaly pressure and advisor debt into a single number. Its one invariant is that a composite never hides a critical: any firing critical pins the score to the bottom band. A score without a drill-down to its components is a vanity metric, and we wrote that into the spec because we know we will be tempted.
Why this matters at 3 a.m.
If you carry a database pager, you have run an automation script that somebody wrote after the last outage, and you have wondered whether it would cause the next one. The ladder is our answer to that. An action climbs only after a measured record on your fleet. The top rung requires a rollback that exists, not one that is described. Destructive operations have no rung at all.
We capped ourselves at rung 2 because rung 3 is a promise, and we only make it where the tests can prove we keep it.