The platform
PostgreSQL performance monitoring that ends in a governed action
DBExplore is one agent, one console, and one policy gate across every Postgres engine your teams run. Here is what each layer does.
Fleet observability
One console for every Postgres you run
DBExplore collects the same signals from managed clouds, serverless platforms, distributed engines and self-hosted clusters, then normalizes them so a replication lag on Aurora and a replica lag on Patroni land on the same chart with the same meaning.
Active session history is sampled continuously, so a two-minute lock storm at 03:14 is still there when you look at 09:00. Query plans are captured by structure, so you see plan changes, not plan noise.
- Active session history, wait events, and period-over-period compare
- Query performance with plan capture, plan-change detection, and regression alerts
- Locks, vacuum and wraparound runway, bloat, WAL, replication and HA topology, and connection poolers
- Config and schema drift against golden baselines, and security posture
- Seven-dimension topology discovery per cluster, re-checked on a schedule, so a changed shape is an event rather than a surprise
Lock:transactionid ████████████████░░░░ 61% IO:DataFileRead ██████░░░░░░░░░░░░░░ 22% CPU ████░░░░░░░░░░░░░░░░ 12% blocking pid 48211 ← 14 waiters · UPDATE orders SET …
Anomaly detection
Hundreds of detection rules, tuned per engine, plus models for what rules cannot see
Aurora noise is not CockroachDB signal. Every detection rule carries the engine it applies to and its own suppression behaviour, so a still-firing anomaly is one event with a duration, not four hundred identical alerts.
Signatures cover the failure modes operators can name. Statistical models cover the ones nobody has written a runbook for yet.
- Curated detection rules across infrastructure, workload, query, schema, and operational categories
- Multivariate anomaly scoring with adaptive baselines, and plan-regression detection that needs both a plan change and a real slowdown
- Deduplication and flap suppression so incidents read as one story, plus capacity forecasting expressed as runway
cluster pg-payments engine postgresql leader changed 3× in 4 min one event, deduped severity CRITICAL repeats suppressed runbook attached → open
AI advisor
Recommendations with the evidence attached
The advisor proposes indexes, query rewrites, and schema-safety changes and shows its work: the plan before, the plan after, the estimated cost delta, and the writes the new index will slow down. A recommendation you cannot audit is a guess with a confident tone.
Natural-language questions about your fleet run through a read-only gate. The model drafts, the gate verifies it reads only, and you confirm before anything executes.
- Index and query advice with what-if analysis and return-on-investment ranking
- Workload-aware suggestions, so write-heavy tables never get advice that hurts them
- Ask your fleet in plain English through a fail-closed, read-only gate
CREATE INDEX CONCURRENTLY ON orders (tenant_id, created_at) what-if seq scan 812 ms → index scan 47 ms (−94%) writes +3.1% on INSERT orders size ~1.9 GB evidence 3 queries · plan diff attached
Action plane
Remediation that dry-runs first and verifies after
Dozens of remediation templates cover the fixes DBAs actually run, from cancelling a runaway query to adding an index concurrently or reloading a pooler.
Every template declares its safety class and reversibility. Every run dry-runs, executes, then probes the live system to prove the change took effect. A no-op is reported as a no-op, never as a success.
- Safety classes from read-only to disruptive, declared per template
- Dry-run, execute, verify-after-act, and pre-computed rollback for reversible actions
- Cooldowns and locking so the same action cannot storm a target or collide with another operator
safety local mutation reversible client retries dry-run pid 48211 running 14m 03s · app=etl-nightly verify live probe after execute cooldown active per target · lock held
The path every fix takes
Nothing executes as free-form SQL.
Every fix is an adapter with a declared safety class, a dry run, a verifier and, where the world allows it, a rollback. There is no path from a model straight to your database.
- dry run
Predict the effect against live state. No mutation. Always available.
- gate
Policy-as-code decides. Deny by default, fail closed.
- approve
A human. Two of them for anything cluster-wide or disruptive.
- execute
The adapter runs. Never free-form SQL, only a declared template.
- verify
Probe the live system. A no-op is reported as a no-op, never as success.
- rollback
Where the action declared reversibility and implemented a rollback.
Policy gate
Fail-closed. Human approval by default. Autonomy is earned.
Before any mutating action runs, a policy engine evaluates it. Deny, timeout, unreachable, malformed response, undefined policy: every outcome that is not an explicit allow is a deny. If the gate is down, nothing mutates.
Tenants start at approve. Disruptive actions require more than one approver and can never be switched to auto. An action graduates to one-click or auto-apply only after a measured precision record, per tenant, per action class.
- Policy-as-code rules per tool, evaluated before the runtime starts
- Autonomy ladder: observe, recommend, one-click, auto-apply, each with stricter preconditions
- Multi-approver rule for disruptive actions, a fleet-wide kill switch, and expiring break-glass access
Add index on orders(tenant_id, created_at)
seq scan → index scan · est. −94% query time · prod-eu-aurora-02
- Policy gate passed
- Fully reversible
- Inside off-peak window
Five ways to get a deny
Anything that is not an explicit allow becomes a no.
There is no ambient-authority path and no bypass flag. If the gate cannot say yes, the answer is no, and the action does not run.
- deny
The policy evaluated the action and refused it.
- timeout
The policy engine did not answer in time.
- unreachable
The policy engine could not be reached at all.
- malformed
The response came back, but not in a shape we can trust.
- undefined
No policy exists yet for this action.
If the gate itself is down, automation is down. Your database is not.
Trust ledger
Every observation, decision, and action, signed
Each tenant has an append-only ledger. Entries are chain-hashed and signed with keys that rotate on a schedule, and periodic snapshots give auditors inclusion proofs without handing them the whole log.
Erasure requests are honoured by destroying the tenant data key, so the chain stays intact and the data is unrecoverable. That is how long retention and a deletion request coexist.
- Signed, chain-hashed entries with periodic snapshots per tenant
- Ledger verification in the console and through the MCP server
- Long-term retention and export in a documented format for your auditors
entries 61 signed · keys rotated on schedule chain ok snapshot root 7c1e…a90f verify 61/61 valid 0 gaps 0 tamper export auditor bundle
Proof, not logs
A log can be edited. A chain cannot.
Every observation, decision and action, by a human or by an agent, lands in a per-tenant, append-only ledger. Each entry carries a hash of its own payload and a hash of the entry before it, then the whole thing is signed. Change one entry and every entry after it stops verifying.
- entry 1
anomaly observed
- payload hash
- …21a0f
- prev hash
- genesis
- signature
- valid
- entry 2
fix proposed
- payload hash
- …28a1f
- prev hash
- …21a0f
- signature
- valid
- entry 3
human approved
- payload hash
- …35a2f
- prev hash
- …28a1f
- signature
- valid
- entry 4
action verified
- payload hash
- …42a3f
- prev hash
- …35a2f
- signature
- valid
Verify any range yourself, from the console or through the MCP server. Hand the export to an auditor without translating anything.
MCP server
Your AI assistants use the same guardrails, never a way around them
DBExplore ships an MCP server so Claude, Cursor, or your own agents can investigate a slow query, read a ledger range, or propose a remediation through exactly the same policy gate a human would hit.
Coverage is declared tool by tool. No hidden tools, no privileged shortcut.
- Fleet, query, plan, anomaly, ledger, advisor, and action tool families
- Authenticated, tenant-scoped, read-only by default
- Every mutating tool call passes the same fail-closed gate and lands in the same ledger
fleet.list_clusters ✓ query.top_regressions ✓ ledger.verify_range ✓ action.propose gate · approve required
Integrations
Fits the on-call stack you already run
Alerts route to Slack, PagerDuty, OpsGenie, Microsoft Teams, Jira, email, and webhooks with escalation and on-call schedules. OpenTelemetry logs ingest natively, and metrics and events export to the dashboards your SREs already watch.
- Slack, PagerDuty, OpsGenie, Teams, Jira, email, and webhooks with escalation policies
- OpenTelemetry ingestion and export to Grafana and Datadog
- Slack-native approvals and an SLO platform with error-budget burn alerts
slack #db-oncall posted 03:14:07 pagerduty EP database-primary triggered jira DBA-2291 created linked to incident approval ✓ via Slack by s.okafor 03:19:52
API and SDKs
An API with dashboards as one option
Everything the console shows is available through the REST API: fleet inventory, anomalies, plans, advisor findings, ledger ranges, and remediation proposals. SDKs for Python, TypeScript, and Go wrap it with typed models, so the pipeline that provisions the rest of your infrastructure can drive DBExplore as code.
- REST API with tenant-scoped, read-only-capable API keys
- Python, TypeScript, and Go SDKs
- A headless mode for building your own portal on top
200 OK · 12 items · tenant-scoped key
{ "signature": "leader-flapping",
"severity": "critical", "runbook": "…" }
sdk from dbexplore import Client The agent team
Four roles, separated on purpose.
One model doing everything is one prompt away from a bad afternoon. DBExplore splits the work the way a database team does, and only the last role can touch anything.
- 01
Observers
Watch the fleet continuously and turn raw telemetry into named, deduplicated conditions. They never propose anything.
Read only
- 02
Analysts
Take a condition and work out why. Correlate across the cluster, the topology and recent change, and produce a root cause with the evidence attached.
Read only
- 03
Advisors
Turn a cause into a specific, costed recommendation, with the plan before, the plan after, and what it will slow down.
Proposes
- 04
Remediators
Carry an approved recommendation through the gate, execute the declared template, verify against live state, and sign the result.
Acts, under policy
Each role is specialised per signal family rather than general-purpose, so the agent reasoning about replication lag is not the same one reasoning about vacuum. A finding has to survive the handoff between roles, which is a cheaper filter than asking one model to check its own work.
Co-learning
It gets better on your fleet, without your data leaving it.
A detection library that ships the same thresholds to everyone is wrong for almost everyone. DBExplore learns from what your team approves, rejects and rolls back, and tunes itself to your fleet.
Every outcome is a label
Approved, rejected, rolled back, ignored. Each one is a signal about whether that recommendation was worth making on your fleet.
Calibrated against your baselines
Thresholds that fire correctly on a busy payments cluster are wrong on a quiet reporting replica. They are learned per cluster, not shipped as one number.
Ranked by what you actually act on
Advice you consistently skip drops down the queue. Advice you consistently take moves up, and eventually becomes a candidate for one-click.
Rolled out in the shade first
Every loop starts shadowed, comparing itself against what actually happened without touching anything. Then canary. Then on, per tenant, by your choice.
Your learning stays yours
The model starts from a general prior so a new cluster is useful on day one, then everything it learns about your fleet stays scoped to your tenant. Training data never crosses a tenant boundary. Nobody else's fleet teaches on your incidents, and yours does not teach on theirs.
Nothing is switched on for you
Each loop is independently flagged per tenant and starts in shadow mode, where it makes predictions and records whether it would have been right, while changing nothing. You promote it when the record justifies it. This is the same discipline the autonomy ladder uses, applied to learning rather than to acting.
See the platform on your own fleet.
Agent install takes well under an hour. Real signals fire within the first day. Tell us about your engines and we will scope a pilot.