Skip to content

Incidents (dw-incidents)

The Incidents agent is the on-call responder for your data platform. Feed it anomaly signals — a metric that moved, a pipeline that failed, an SLA that slipped — and it classifies the incident into one of six types (schema change, source delay, resource exhaustion, code regression, infrastructure, quality degradation), assigns a severity, and suggests remediation. It works the way a good SRE does: timeline first, then competing hypotheses, then evidence, then escalation.

For root cause, it doesn’t guess from the symptom. It traverses the lineage graph upstream, queries execution logs, and cross-references incident history to build a causal chain with confidence scores — distinguishing the symptom (a missing row count) from the cause (a late-arriving upstream batch, a schema change that broke a join).

  • Incident classification. diagnose_incident takes raw anomaly signals and returns an incident type, severity, and suggested remediation actions.
  • Graph-based root cause analysis. get_root_cause traverses lineage up to five or more hops upstream and returns a causal chain with confidence scores, not a single guess.
  • Playbook remediation with a confidence gate. remediate runs known playbooks — restart task, scale compute, apply schema migration, switch backup source, backfill data. Automatic execution requires diagnosis confidence above 0.95; anything novel generates a diagnosis report and routes to a human for approval.
  • Dry-run mode. Simulate a remediation before executing it to see exactly what would happen.
  • Incident memory. get_incident_history uses vector similarity to find past incidents like the current one, so the third occurrence of a pattern is diagnosed faster than the first.
  • Structured incident communication. Output leads with severity, phase, and impact surface — which assets, which downstream consumers, since when — in scannable form.

“Row counts on analytics.daily_orders dropped 40% overnight — diagnose it.”

“What’s the root cause of this incident? Walk the lineage upstream and show me the causal chain.”

“Have we seen an incident like this before? Show me similar past incidents.”

“Dry-run the backfill playbook for this incident before we execute anything.”

  • Warehouses and lakehouses — Snowflake, BigQuery, Databricks for the signals and the affected assets.
  • Orchestration — Airflow and friends, for pipeline run context and restart playbooks.
  • Alerting — PagerDuty, Slack, Microsoft Teams, OpsGenie for paging and resolution.

See the connector catalog for setup.

The agent starts in 🟡 Evaluation on built-in sample data: a realistic set of incidents, metrics, and lineage you can diagnose end to end before any credential exists. It earns 🟢 Connected per system through a passing live test. See Verify your setup.

  • In 🟡 Evaluation, diagnosis and remediation run against the sample estate — no real system is touched until connections are verified.
  • Auto-remediation is deliberately narrow: only known patterns above the 0.95 confidence threshold execute without a human. Everything else stops at a diagnosis report and an approval request. That’s a floor, not a temporary restriction.
  • Remediation playbooks are a fixed set; incidents outside them route to a human with the evidence gathered so far rather than an improvised fix.
  • Root cause quality depends on lineage coverage. If an upstream system isn’t connected, the causal chain stops at the boundary and says so.