Skip to content

Anomaly Detection

Monitor ships a small, built-in analyzer (internal/analyzer) that watches the same metric stream the TUI and CLI render, and raises alerts when something looks wrong. You don't configure it to exist — it runs on every watch tick — you only configure which thresholds matter to you.

Where alerts come from

The collector publishes one Event per tick. The analyzer observes each event and returns zero or more collector.Alert objects. monitor watch is the surface that streams them: every alert becomes an alert line in the NDJSON stream, and the same alert can fan out to a desktop notification, a webhook, and an incident stash.

collector (tick) ──► analyzer ──► alert(s)

        ┌──────────────────────────┼──────────────────────────┐
        ▼              ▼            ▼                          ▼
   NDJSON line     --notify     --webhook POST          --stash (fcheap)

Every alert carries the same shape regardless of which rule fired:

json
{
  "rule": "cpu_spike",
  "severity": "warning",
  "pid": 1133,
  "process": "node",
  "detail": "cpu spike",
  "diagnosis": {
    "summary": "node pinned a core for 45s while RSS stayed flat — consistent with a hot loop",
    "evidence": ["cpu 150% for 45s", "rss flat"],
    "confidence": "medium",
    "next_actions": ["monitor profile 1133 --type cpu", "monitor investigate 1133"]
  }
}

pid and process are present only on per-process rules; system-wide rules omit them. diagnosis is attached whenever the analyzer's cross-signal rule table can interpret the pattern — see Diagnosis below.

The rules

Five rules always run, and two more light up when you configure them.

Always-on rules

RuleWhat it detectsDefault threshold
cpu_spikeCurrent process CPU is at least 3× the median of its prior samples. Three baseline samples and an absolute floor prevent idle jitter from firing.factor 3.0, floor 50%
rss_growthResident memory is climbing steadily — a regression over RSS history, normalized by elapsed wall-clock time, with slope > ~50 KB/s and R² > 0.7.50_000 bytes/second
disk_fillAny mounted partition whose usage is at or above the threshold.90%
swap_pressureSwap usage at or above a fraction of swap total — real memory pressure, not just high RAM use.50% of swap total
zombie_processThe OS reports a process in zombie (Z) state, waiting for its parent to reap it. Unsupported status telemetry stays quiet.status Z

Per-process CPU is based on two cumulative user+system CPU counter samples, divided by elapsed wall time. 100% means one fully occupied core and a multithreaded process may exceed it. The first observation of a process has no delta; PID reuse, a backwards counter, or a non-positive interval is also reported unavailable. Those samples do not pretend to be an observed zero and cannot seed a misleading cpu_spike baseline. Full-process collection uses a minimum 100ms cadence.

cpu_spike, rss_growth, and zombie_process are per-process (the alert carries the PID); disk_fill fires once per offending partition; swap_pressure and the threshold rules are system-wide.

Configurable threshold rules

The cpu_alert_threshold and memory_alert_threshold settings in config.json are dormant until you set them above 0. When set, a ThresholdRule fires:

  • cpu_threshold — overall CPU usage >= the configured percentage.
  • mem_threshold — overall memory usage >= the configured percentage.

A threshold of 0 disables that check (the default). Set them in the Settings tab or by editing the config file directly:

json
{
  "cpu_alert_threshold": 90,
  "memory_alert_threshold": 85
}

These are percentage-point overall-usage alerts, distinct from the per-process cpu_spike / rss_growth rules.

Diagnosis

Raw alerts say what crossed a line; the diagnosis says why it matters and what to do next. Every diagnosis has four fields:

FieldMeaning
summaryOne plain-language sentence, e.g. "RSS grew 42%/10min while CPU stayed flat — consistent with a memory leak (slope 3.2MB/min, R²=0.94)".
evidenceThe numbers behind the claim (regression slope, R², durations, sample counts) — never discarded, so an agent can double-check the reasoning.
confidencelow, medium, or high, graded from the strength of the evidence: regression fit (R², sample count) for a trend, the level for a flat-but-high signal, or the reversal count for a sawtooth.
next_actionsAt most two concrete next steps — a monitor CLI command or MCP tool invocation.

The analyzer (internal/analyzer/diagnosis.go) correlates RSS and CPU trend classes over the per-PID history it already holds, using a small ordered rule table (most specific pattern first):

PatternInterpretation
RSS sawtooth (reversals, weak linear fit)GC pressure / allocation churn
RSS rising and CPU risingload (demand, not a leak)
RSS rising, CPU flat/falling/highmemory leak
RSS flat/falling, CPU high or risingspin / hot loop

Everything else — nothing trending, too few samples — produces no diagnosis; a quiet process stays quiet.

The same Diagnosis shape appears on watch alerts (this page), incident bundles, --notify / --webhook sinks, monitor diff verdicts (see diff), and the monitor_analyze MCP tool (see MCP Server) — one interpretation layer, every surface.

Severity

Every alert the current rules emit is severity: "warning". The field is part of the alert contract so downstream sinks (webhooks, notifications, stashes) can triage consistently as new severities are added.

Alert sinks

Alerts never block the watch loop. Side-effecting sinks share a bounded worker budget (--delivery-limit, default 8); when all slots are busy, Monitor drops the new delivery with an explicit stderr message instead of accumulating unbounded goroutines. See watch for the full flag reference; in summary:

SinkFlagBehavior
NDJSON stream--jsonAlways on with --json; one alert line per alert.
Desktop notification--notifyosascript (macOS) / notify-send (Linux); the message leads with the diagnosis summary (falling back to the terse rule detail when there's no diagnosis) and appends the confidence level; failures log to stderr.
Webhook--webhook <url>POSTs the raw collector.Alert JSON (with diagnosis alongside it when the analyzer attached one); non-2xx logged to stderr.
Incident stash--stashCaptures an incident bundle to fcheap in the background; emits a stash line with the outcome.

Sinks are additive: --notify, --webhook, and --stash all fire alongside the NDJSON stream, not instead of it.

Capturing incidents on alert

Pair the analyzer with --stash to keep a forensic bundle for every alert. Each capture runs in a background goroutine and never stalls the stream; a stash failure surfaces in a stash_error field rather than aborting watch. See Ecosystem Integration for how fcheap stashes work and what happens when fcheap isn't installed.

bash
monitor watch --json --stash --stash-ttl 24h

See also

  • watch — the flags and NDJSON output that stream these alerts.
  • Configuration — the cpu_alert_threshold and memory_alert_threshold settings.
  • Ecosystem Integration — fcheap incident stashes on alert.
  • Architecture — where the analyzer sits in the data flow.

Released under the MIT License.