What our anomaly engine actually runs
An isolation forest with a fixed contamination rate, density clustering with a sample floor, a threshold-ladder classifier, and heuristics that name a cause.
Every database observability vendor has a slide with the words “ML-powered anomaly detection” on it. We have that slide too. This post is the version that says what each method is for, what it is fed, and where we chose to stop it, because that is where the honesty lives.
Isolation Forest, scoped on purpose
The multivariate detector trains on a rolling window of the host and connection signals that move together when a database is in trouble and independently when it is not. If the window is too thin it does nothing, on purpose: a model that scores on a handful of points is a random number generator with a confident name. Contamination is fixed rather than learned, which means the model expects a small, known fraction of any window to be unusual and does not drift toward calling everything normal on a bad week. Only the latest point is scored, and a hit is emitted as its own metric with a severity.
Missing points are forward-filled, so a collector that stalls looks, to the model, like a flat line rather than a gap. That is deliberate. A stalled collector is its own anomaly signature and fires separately.
There is a second forest for I/O timing, with a lower contamination setting because I/O spikes are rare but expensive, and its raw score is scaled to a severity so it can be ranked next to rule-based findings.
Density clustering for regressions, with a floor
Latency distributions for a query are often bimodal. Cache hit, cache miss. A simple average moves when the mix moves, and that is not a regression. So the regression detector clusters. It needs a minimum number of samples before it will say anything. The neighbourhood size scales with the median so that a query that normally takes a millisecond and one that takes a second are judged by the same standard. The baseline is the dominant historical cluster; stragglers are ignored. The current window is a regression if it sits well above that baseline, an improvement if it sits well below, and otherwise a new mode. If the current batch is all stragglers, a plain median comparison takes over at a stricter multiple.
For distributed engines the regression multiple relaxes, because cross-node coordination and rebalancing make modest shifts routine and the tighter rule produced false positives. A sub-threshold shift on those engines is still surfaced as a new mode. It is never hidden.
The workload classifier is a ladder, on purpose
Workload fingerprinting classifies a database as read-heavy, write-heavy, analytical or mixed, so the advisor does not propose an index on a table taking thousands of inserts a second. It is a threshold ladder over read ratio and scan mix, not a clustering model, and each rung carries a confidence.
We tried clustering first. A ladder is explainable in one sentence to a DBA who disagrees with it, and a classifier the DBA cannot argue with is a classifier the DBA will turn off.
Disk root cause, four heuristics
When combined disk latency crosses a saturation threshold, the disk analyser adds four heuristics into a capped confidence. A high share of forced checkpoints points at WAL sizing. Too many concurrent autovacuum workers points at a vacuum storm. High read latency with healthy writes points at a cold cache or a missing index. The write-side mirror points at the opposite. Each heuristic names its cause and its fix in the finding.
These are heuristics and they say so. They are right often enough to be the first thing an operator checks.
The concurrency predictor
Waiting sessions are tracked with a first and second derivative over a short window. A high absolute count is high. A moderate count that is growing is medium, and if the growth is accelerating, high. Fast growth is critical regardless of level. The trade-off we chose is no hysteresis: it reacts within a few intervals and it can flap at low sample counts. The suppression layer below absorbs that, and we preferred an early warning that repeats to a late one that is tidy.
What suppresses the noise
None of the above would be usable without the layer underneath. Every anomaly gets a stable fingerprint from tenant, cluster, signature and the dedup keys the signature declares. The warehouse is ordered on that fingerprint so the latest state per anomaly is a cheap lookup, and that is the basis of the alert state machine.
Flap suppression sits on top. Each signature can declare a minimum interval between emissions. Suppressed firings still update the dedup record, so the console shows the anomaly as still firing with its original start time; we just stop re-emitting identical rows every few seconds. On restart the tracker seeds itself from recent history so a fresh process does not unleash every suppressed event at once.
Feedback scales a signature’s confidence within a bounded range. A broken feedback callback defaults to no change rather than silently muting a detector, because a detector that goes quiet is worse than one that is occasionally wrong.
ML and rules are allowed to disagree
We run both, and we wrote down when to trust which. Signatures win for anything an operator can express as a WHERE clause on state: a replication slot with status lost, a leader that changed three times in four minutes, a quorum that is gone. They carry a runbook and they are stable across restarts. Models win for continuous numerics with failure modes nobody has named yet. Model output is never an action-plane target, because it is too noisy to auto-execute, and when both fire on the same incident the signature is the one that pages, because it knows what to do next.
What we measure
Two of the conditions above have reference pages of their own, written for anyone who would rather run the queries themselves: transaction ID wraparound and lock contention and blocking trees.
Every detector on this page is held to gates rather than to a slide: a high per-action precision record over a sustained period before an action reaches one-click, a stricter one before auto-apply, and calibrated confidence before we trust a probability the model reports. Design partners are producing the first fleet-scale numbers now, and when we publish them they will come with the same level of explanation this post has, because a precision figure without its method is a slide, not a result.