Skip to content
dbexplore

The platform

PostgreSQL performance monitoring that ends in a governed action

DBExplore is one agent, one console, and one policy gate across every Postgres engine your teams run. Here is what each layer does.

Fleet observability

One console for every Postgres you run

DBExplore collects the same signals from managed clouds, serverless platforms, distributed engines and self-hosted clusters, then normalizes them so a replication lag on Aurora and a replica lag on Patroni land on the same chart with the same meaning.

Active session history is sampled continuously, so a two-minute lock storm at 03:14 is still there when you look at 09:00. Query plans are captured by structure, so you see plan changes, not plan noise.

  • Active session history, wait events, and period-over-period compare
  • Query performance with plan capture, plan-change detection, and regression alerts
  • Locks, vacuum and wraparound runway, bloat, WAL, replication and HA topology, and connection poolers
  • Config and schema drift against golden baselines, and security posture
  • Seven-dimension topology discovery per cluster, re-checked on a schedule, so a changed shape is an event rather than a surprise
ash · prod-eu-aurora-02 · last 5 minillustrative

Anomaly detection

Hundreds of detection rules, tuned per engine, plus models for what rules cannot see

Aurora noise is not CockroachDB signal. Every detection rule carries the engine it applies to and its own suppression behaviour, so a still-firing anomaly is one event with a duration, not four hundred identical alerts.

Signatures cover the failure modes operators can name. Statistical models cover the ones nobody has written a runbook for yet.

  • Curated detection rules across infrastructure, workload, query, schema, and operational categories
  • Multivariate anomaly scoring with adaptive baselines, and plan-regression detection that needs both a plan change and a real slowdown
  • Deduplication and flap suppression so incidents read as one story, plus capacity forecasting expressed as runway
anomaly · leader flappingillustrative

AI advisor

Recommendations with the evidence attached

The advisor proposes indexes, query rewrites, and schema-safety changes and shows its work: the plan before, the plan after, the estimated cost delta, and the writes the new index will slow down. A recommendation you cannot audit is a guess with a confident tone.

Natural-language questions about your fleet run through a read-only gate. The model drafts, the gate verifies it reads only, and you confirm before anything executes.

  • Index and query advice with what-if analysis and return-on-investment ranking
  • Workload-aware suggestions, so write-heavy tables never get advice that hurts them
  • Ask your fleet in plain English through a fail-closed, read-only gate
advisor · index candidateillustrative

Action plane

Remediation that dry-runs first and verifies after

Dozens of remediation templates cover the fixes DBAs actually run, from cancelling a runaway query to adding an index concurrently or reloading a pooler.

Every template declares its safety class and reversibility. Every run dry-runs, executes, then probes the live system to prove the change took effect. A no-op is reported as a no-op, never as a success.

  • Safety classes from read-only to disruptive, declared per template
  • Dry-run, execute, verify-after-act, and pre-computed rollback for reversible actions
  • Cooldowns and locking so the same action cannot storm a target or collide with another operator
action · cancel long queryillustrative

The path every fix takes

Nothing executes as free-form SQL.

Every fix is an adapter with a declared safety class, a dry run, a verifier and, where the world allows it, a rollback. There is no path from a model straight to your database.

  1. dry run

    Predict the effect against live state. No mutation. Always available.

  2. gate

    Policy-as-code decides. Deny by default, fail closed.

  3. approve

    A human. Two of them for anything cluster-wide or disruptive.

  4. execute

    The adapter runs. Never free-form SQL, only a declared template.

  5. verify

    Probe the live system. A no-op is reported as a no-op, never as success.

  6. rollback

    Where the action declared reversibility and implemented a rollback.

Destructive operations have no adapter at all. Dropping a table, truncating one, or forcing a failover cannot be requested, by a person or by a model.

Policy gate

Fail-closed. Human approval by default. Autonomy is earned.

Before any mutating action runs, a policy engine evaluates it. Deny, timeout, unreachable, malformed response, undefined policy: every outcome that is not an explicit allow is a deny. If the gate is down, nothing mutates.

Tenants start at approve. Disruptive actions require more than one approver and can never be switched to auto. An action graduates to one-click or auto-apply only after a measured precision record, per tenant, per action class.

  • Policy-as-code rules per tool, evaluated before the runtime starts
  • Autonomy ladder: observe, recommend, one-click, auto-apply, each with stricter preconditions
  • Multi-approver rule for disruptive actions, a fleet-wide kill switch, and expiring break-glass access

Five ways to get a deny

Anything that is not an explicit allow becomes a no.

There is no ambient-authority path and no bypass flag. If the gate cannot say yes, the answer is no, and the action does not run.

  • deny

    The policy evaluated the action and refused it.

  • timeout

    The policy engine did not answer in time.

  • unreachable

    The policy engine could not be reached at all.

  • malformed

    The response came back, but not in a shape we can trust.

  • undefined

    No policy exists yet for this action.

If the gate itself is down, automation is down. Your database is not.

Trust ledger

Every observation, decision, and action, signed

Each tenant has an append-only ledger. Entries are chain-hashed and signed with keys that rotate on a schedule, and periodic snapshots give auditors inclusion proofs without handing them the whole log.

Erasure requests are honoured by destroying the tenant data key, so the chain stays intact and the data is unrecoverable. That is how long retention and a deletion request coexist.

  • Signed, chain-hashed entries with periodic snapshots per tenant
  • Ledger verification in the console and through the MCP server
  • Long-term retention and export in a documented format for your auditors
ledger · range 41,200–41,260illustrative

Proof, not logs

A log can be edited. A chain cannot.

Every observation, decision and action, by a human or by an agent, lands in a per-tenant, append-only ledger. Each entry carries a hash of its own payload and a hash of the entry before it, then the whole thing is signed. Change one entry and every entry after it stops verifying.

  1. entry 1

    anomaly observed

    payload hash
    …21a0f
    prev hash
    genesis
    signature
    valid
  2. entry 2

    fix proposed

    payload hash
    …28a1f
    prev hash
    …21a0f
    signature
    valid
  3. entry 3

    human approved

    payload hash
    …35a2f
    prev hash
    …28a1f
    signature
    valid
  4. entry 4

    action verified

    payload hash
    …42a3f
    prev hash
    …35a2f
    signature
    valid

Verify any range yourself, from the console or through the MCP server. Hand the export to an auditor without translating anything.

MCP server

Your AI assistants use the same guardrails, never a way around them

DBExplore ships an MCP server so Claude, Cursor, or your own agents can investigate a slow query, read a ledger range, or propose a remediation through exactly the same policy gate a human would hit.

Coverage is declared tool by tool. No hidden tools, no privileged shortcut.

  • Fleet, query, plan, anomaly, ledger, advisor, and action tool families
  • Authenticated, tenant-scoped, read-only by default
  • Every mutating tool call passes the same fail-closed gate and lands in the same ledger
mcp · tools/listillustrative

Integrations

Fits the on-call stack you already run

Alerts route to Slack, PagerDuty, OpsGenie, Microsoft Teams, Jira, email, and webhooks with escalation and on-call schedules. OpenTelemetry logs ingest natively, and metrics and events export to the dashboards your SREs already watch.

  • Slack, PagerDuty, OpsGenie, Teams, Jira, email, and webhooks with escalation policies
  • OpenTelemetry ingestion and export to Grafana and Datadog
  • Slack-native approvals and an SLO platform with error-budget burn alerts
route · sev2 · pg-paymentsillustrative

API and SDKs

An API with dashboards as one option

Everything the console shows is available through the REST API: fleet inventory, anomalies, plans, advisor findings, ledger ranges, and remediation proposals. SDKs for Python, TypeScript, and Go wrap it with typed models, so the pipeline that provisions the rest of your infrastructure can drive DBExplore as code.

  • REST API with tenant-scoped, read-only-capable API keys
  • Python, TypeScript, and Go SDKs
  • A headless mode for building your own portal on top
GET /clusters/pg-payments/anomaliesillustrative

The agent team

Four roles, separated on purpose.

One model doing everything is one prompt away from a bad afternoon. DBExplore splits the work the way a database team does, and only the last role can touch anything.

  1. 01

    Observers

    Watch the fleet continuously and turn raw telemetry into named, deduplicated conditions. They never propose anything.

    Read only

  2. 02

    Analysts

    Take a condition and work out why. Correlate across the cluster, the topology and recent change, and produce a root cause with the evidence attached.

    Read only

  3. 03

    Advisors

    Turn a cause into a specific, costed recommendation, with the plan before, the plan after, and what it will slow down.

    Proposes

  4. 04

    Remediators

    Carry an approved recommendation through the gate, execute the declared template, verify against live state, and sign the result.

    Acts, under policy

Each role is specialised per signal family rather than general-purpose, so the agent reasoning about replication lag is not the same one reasoning about vacuum. A finding has to survive the handoff between roles, which is a cheaper filter than asking one model to check its own work.

Co-learning

It gets better on your fleet, without your data leaving it.

A detection library that ships the same thresholds to everyone is wrong for almost everyone. DBExplore learns from what your team approves, rejects and rolls back, and tunes itself to your fleet.

Every outcome is a label

Approved, rejected, rolled back, ignored. Each one is a signal about whether that recommendation was worth making on your fleet.

Calibrated against your baselines

Thresholds that fire correctly on a busy payments cluster are wrong on a quiet reporting replica. They are learned per cluster, not shipped as one number.

Ranked by what you actually act on

Advice you consistently skip drops down the queue. Advice you consistently take moves up, and eventually becomes a candidate for one-click.

Rolled out in the shade first

Every loop starts shadowed, comparing itself against what actually happened without touching anything. Then canary. Then on, per tenant, by your choice.

Your learning stays yours

The model starts from a general prior so a new cluster is useful on day one, then everything it learns about your fleet stays scoped to your tenant. Training data never crosses a tenant boundary. Nobody else's fleet teaches on your incidents, and yours does not teach on theirs.

Nothing is switched on for you

Each loop is independently flagged per tenant and starts in shadow mode, where it makes predictions and records whether it would have been right, while changing nothing. You promote it when the record justifies it. This is the same discipline the autonomy ladder uses, applied to learning rather than to acting.

See the platform on your own fleet.

Agent install takes well under an hour. Real signals fire within the first day. Tell us about your engines and we will scope a pilot.