# Data Workers Documentation > Onboarding and technical documentation for Data Workers — the autonomous agentic > data platform by Datapokopia: Agent Swarm, Spellbook Data Catalog (research preview), > and Data Context Wizard. Built for pilot, Scale, and Enterprise customers. > Full corpus: https://docs.dataworkers.io/llms-full.txt ## Pages - [Connect Claude Code](https://docs.dataworkers.io/connect/claude-code/): Install Data Workers in Claude Code — plugin, local MCP, or hosted endpoint. - [Connect Codex CLI](https://docs.dataworkers.io/connect/codex/): Install Data Workers in OpenAI Codex CLI via config.toml — local or hosted. - [Connect Cursor](https://docs.dataworkers.io/connect/cursor/): Install Data Workers in Cursor — one-click deeplink or manual mcp.json. - [Connect OpenCode](https://docs.dataworkers.io/connect/opencode/): Install Data Workers in OpenCode via opencode.json — local or hosted. - [Other MCP clients](https://docs.dataworkers.io/connect/other-clients/): Data Workers in Claude Desktop, VS Code, Windsurf, Cline, Continue, ChatGPT, Gemini CLI, and any generic MCP client. - [Atlan, Alation, or Collibra?](https://docs.dataworkers.io/connectors/atlan-alation-collibra/): The straight answer on the three big commercial catalogs — no native connector today, and the three ways teams connect anyway. - [Connect BigQuery](https://docs.dataworkers.io/connectors/bigquery/): Connect Data Workers agents to BigQuery with a least-privilege service account, two environment variables, and a live verification test. - [Connect Databricks](https://docs.dataworkers.io/connectors/databricks/): Connect Data Workers agents to Databricks Unity Catalog with a least-privilege service principal, two environment variables, and a live verification test. - [Connect DataHub](https://docs.dataworkers.io/connectors/datahub/): Connect Data Workers agents to your DataHub instance with a read-scoped GMS token, two environment variables, and a live verification test. - [Connect dbt](https://docs.dataworkers.io/connectors/dbt/): Connect Data Workers agents to dbt Cloud with a read-only API token, or to a local dbt Core project, and verify with a live test. - [Supported connectors](https://docs.dataworkers.io/connectors/): The 50+ systems Data Workers connects to — warehouses, catalogs, orchestrators, quality, BI, identity — and what's not supported yet. - [Least-privilege access](https://docs.dataworkers.io/connectors/least-privilege/): The minimum grants Data Workers agents need per system, and the rules that keep write access safe. - [Connect OpenMetadata](https://docs.dataworkers.io/connectors/openmetadata/): Connect Data Workers agents to your OpenMetadata instance with a read-scoped bot JWT, two environment variables, and a live verification test. - [Connect Snowflake](https://docs.dataworkers.io/connectors/snowflake/): Connect Data Workers agents to Snowflake with a least-privilege role, five environment variables, and a live verification test. - [Docs for AI agents (MCP)](https://docs.dataworkers.io/help/docs-via-mcp/): Consume these docs from Claude or any agent — llms.txt, llms-full.txt, and the Data Workers docs MCP server with search, read, and feedback tools. - [FAQ](https://docs.dataworkers.io/help/faq/): Frequently asked questions about Data Workers — pricing, security, maturity, connectors, and how it compares. - [Report a bug or request a feature](https://docs.dataworkers.io/help/feedback/): File a bug report, feature request, or incident with Data Workers — as a human through the form, or programmatically from Claude or any agent. - [Troubleshooting](https://docs.dataworkers.io/help/troubleshooting/): Common Data Workers failure modes and their fixes — install, connection, verification, and console issues. - [Data Workers documentation](https://docs.dataworkers.io/): Product documentation for Data Workers — the Agent Swarm, Spellbook Data Catalog, and Data Context Wizard. - [Data Workers Agent Swarm](https://docs.dataworkers.io/products/agent-swarm/): The governed fleet of specialist data agents — what each agent does, how the swarm is structured, and how to run it. - [Catalog & Context (dw-context-catalog)](https://docs.dataworkers.io/products/agents/dw-context-catalog/): Hybrid catalog search, lineage traversal, and one-call asset context — the home of the governed context graph. - [FinOps & Cost (dw-cost)](https://docs.dataworkers.io/products/agents/dw-cost/): Profile warehouse usage, find unused data, estimate savings with real cost math, and get tiered archival recommendations that never auto-execute. - [Data Access & Governance (dw-governance)](https://docs.dataworkers.io/products/agents/dw-governance/): Policy checks, least-privilege access provisioning, column-level PII scanning, and audit reports with a full evidence chain. - [Incidents (dw-incidents)](https://docs.dataworkers.io/products/agents/dw-incidents/): Diagnose data incidents, trace root cause through the lineage graph, and run remediation playbooks — with humans in the loop for anything novel. - [Insights (dw-insights)](https://docs.dataworkers.io/products/agents/dw-insights/): Natural-language analytics against a governed warehouse — with pre-registered analysis plans, reproduced numbers, and anomalies root-caused before they become conclusions. - [Migration (dw-migration)](https://docs.dataworkers.io/products/agents/dw-migration/): Oracle, Teradata, and Redshift to Snowflake migration with a self-correcting translate → judge → parity loop and a completion gate that fails closed. - [MLOps & Models (dw-ml)](https://docs.dataworkers.io/products/agents/dw-ml/): Experiment tracking, a staged model registry, feature pipelines, drift detection, explainability, and A/B testing — the ML lifecycle as an agent. - [Pipelines & Ingestion (dw-pipelines)](https://docs.dataworkers.io/products/agents/dw-pipelines/): Turn a natural-language description into a validated, deployable data pipeline — with templates, sandbox validation, and Airflow deployment. - [Quality (dw-quality)](https://docs.dataworkers.io/products/agents/dw-quality/): Five-dimension quality scoring, statistical anomaly detection against 14-day baselines, and SLAs that alert with evidence instead of noise. - [Schema (dw-schema)](https://docs.dataworkers.io/products/agents/dw-schema/): Detect schema changes, classify them as breaking or not, assess blast radius through lineage, and generate migrations with rollback built in. - [Data & Cloud Security (dw-security)](https://docs.dataworkers.io/products/agents/dw-security/): DSPM and CSPM in one agent — discover and classify sensitive data, scan cloud posture, and route ranked findings to the agent that can fix them. - [Streaming (dw-streaming)](https://docs.dataworkers.io/products/agents/dw-streaming/): Kafka Connect configuration generation, consumer-lag monitoring, stream health, and tuning recommendations that are always human-applied. - [Usage Intelligence (dw-usage-intelligence)](https://docs.dataworkers.io/products/agents/dw-usage-intelligence/): Zero-LLM analytics on how practitioners actually use the platform — adoption, sessions, workflow patterns, heatmaps, and a hash-chained activity log. - [Conductor](https://docs.dataworkers.io/products/conductor/): The always-on operations loop for your data estate — propose-first, approval-gated, in preview. - [Data Workers Data Context Wizard](https://docs.dataworkers.io/products/data-context-wizard/): The governed, federated knowledge graph your agents read, reason over, and write back to — context your agents can trust. - [Deployment options](https://docs.dataworkers.io/products/deployment/): Self-hosting the Data Workers agent swarm — npx, source, Docker Compose, Kubernetes, and air-gapped installs. - [Integrations & SDKs](https://docs.dataworkers.io/products/integrations/): ChatGPT apps, agent-to-agent protocol, framework adapters, the Python SDK, and the catalog provider SDK. - [Spellbook Data Catalog](https://docs.dataworkers.io/products/spellbook-data-catalog/): The browser control plane over the agent fleet — currently in research preview. - [Warming up your knowledge graph](https://docs.dataworkers.io/products/warming-your-graph/): How the Data Context Wizard goes from a cold start to trusted, authoritative context — what's automated and what only you can provide. ## For agents - File bugs/features/incidents: POST https://docs.dataworkers.io/api/feedback (see /help/feedback/) - Docs MCP server: npx -y @dataworkers/docs-mcp (tools: search_docs, read_doc, submit_feedback) - Preserve maturity labels (research preview / not supported) when quoting these docs. ============================== PAGE: Connect Claude Code URL: https://docs.dataworkers.io/connect/claude-code/ DESCRIPTION: Install Data Workers in Claude Code — plugin, local MCP, or hosted endpoint. ============================== Three ways to wire Data Workers into Claude Code, from most to least recommended. ## Option 1 — The plugin (recommended) ```bash claude plugin marketplace add Datapokopia/data-workers claude plugin install data-workers ``` Verify inside Claude Code with `/mcp` — you should see `dataworkers` connected with its tools listed. ## Option 2 — Local MCP server One line: ```bash claude mcp add data-workers -- npx -y dw-claw ``` Or in a project's `.mcp.json`: ```json { "mcpServers": { "dataworkers": { "type": "stdio", "command": "npx", "args": ["-y", "dw-claw"], "env": { "DW_TELEMETRY_OPT_IN": "false", "DW_MCP_SOURCE": "claude-code" } } } } ``` Local stdio installs send no telemetry unless you opt in. ## Option 3 — Hosted endpoint (Scale and Enterprise) If your workspace uses the hosted endpoint, add it to `~/.claude.json`: ```json { "mcpServers": { "dataworkers": { "type": "http", "url": "https://mcp.dataworkers.io/v1", "headers": { "Authorization": "Bearer ", "X-Mcp-Source": "claude-code" } } } } ``` The hosted endpoint exposes 28 governed tools spanning the fleet. ## First prompts - *"Scan the customer schema for PII."* - *"What's the status of my data connectors?"* — the live-verification report; see [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/) - *"Why did the orders table row count drop yesterday?"* ## Next Connect a real system: [Connect your data systems](/connectors/). Problems installing: [Troubleshooting](/help/troubleshooting/), or run `npx -y @dataworkers/dw-claw@latest doctor`. ============================== PAGE: Connect Codex CLI URL: https://docs.dataworkers.io/connect/codex/ DESCRIPTION: Install Data Workers in OpenAI Codex CLI via config.toml — local or hosted. ============================== Add Data Workers to `~/.codex/config.toml`. ## Local install ```toml [mcp_servers.dataworkers] command = "npx" args = ["-y", "dw-claw"] env_vars = ["DW_LOG_LEVEL"] startup_timeout_sec = 30 tool_timeout_sec = 90 default_tools_approval_mode = "prompt" ``` The `prompt` approval mode means Codex asks before each Data Workers tool runs. Once you trust a read-only tool, you can auto-approve it individually: ```toml [mcp_servers.dataworkers.tools.lineage_read] approval_mode = "auto" ``` ## Hosted (Scale and Enterprise) ```toml [mcp_servers.dataworkers] url = "https://mcp.dataworkers.io/v1" bearer_token_env_var = "DW_API_KEY" ``` Export `DW_API_KEY` in the shell where you run `codex`. ## Verify Start `codex` and list MCP servers — `dataworkers` should appear with its tool set. Then: > Test the connection to my Snowflake catalog. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/) for what the three states mean. ============================== PAGE: Connect Cursor URL: https://docs.dataworkers.io/connect/cursor/ DESCRIPTION: Install Data Workers in Cursor — one-click deeplink or manual mcp.json. ============================== ## One-click install Visit [app.dataworkers.io/connect/cursor](https://app.dataworkers.io/connect/cursor) and use the install button. The deeplink carries a **15-minute, single-use** install token — a deliberate mitigation against deeplink-hijacking attacks, so if you wait too long, just refresh the page for a fresh link. ## Manual install Local install, in `.cursor/mcp.json` at your project root (or `~/.cursor/mcp.json` globally): ```json { "mcpServers": { "dataworkers": { "command": "npx", "args": ["-y", "dw-claw", "all"] } } } ``` Hosted (Scale and Enterprise): ```json { "mcpServers": { "dataworkers": { "url": "https://mcp.dataworkers.io/v1", "headers": { "Authorization": "Bearer ", "X-Mcp-Source": "cursor" } } } } ``` Restart Cursor, then check Settings → MCP shows `dataworkers` with a green status. :::note Cursor Background Agent support is upcoming, not shipped — agents run in your interactive session today. ::: ## First prompts - *"Scan the customer schema for PII."* - *"Which dashboards break if I rename `customer_id`?"* Next: [connect a real system](/connectors/), then [verify it](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ============================== PAGE: Connect OpenCode URL: https://docs.dataworkers.io/connect/opencode/ DESCRIPTION: Install Data Workers in OpenCode via opencode.json — local or hosted. ============================== Add Data Workers to `opencode.json` in your project (or `~/.config/opencode/opencode.json` globally). ## Local install ```jsonc { "$schema": "https://opencode.ai/config.json", "mcp": { "data-workers": { "type": "local", "command": ["npx", "-y", "dw-claw"], "environment": { "DW_MODE": "local", "DW_TELEMETRY": "off" }, "enabled": true, "timeout": 30000 } } } ``` ## Hosted (Scale and Enterprise) ```jsonc { "mcp": { "data-workers": { "type": "remote", "url": "https://mcp.dataworkers.io/v1", "headers": { "Authorization": "Bearer {env:DATAWORKERS_API_KEY}" }, "enabled": true } } } ``` :::note A deeper OpenCode plugin (beyond the MCP server) is on the roadmap. The MCP path above is fully supported today. ::: ## Verify Restart OpenCode, confirm `data-workers` appears in the MCP list, then ask: > What's the status of my data connectors? See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ============================== PAGE: Other MCP clients URL: https://docs.dataworkers.io/connect/other-clients/ DESCRIPTION: Data Workers in Claude Desktop, VS Code, Windsurf, Cline, Continue, ChatGPT, Gemini CLI, and any generic MCP client. ============================== Data Workers agents are standard MCP servers, so any MCP-capable client can run them. The four first-class harnesses have dedicated pages — [Claude Code](/connect/claude-code/), [Cursor](/connect/cursor/), [Codex CLI](/connect/codex/), [OpenCode](/connect/opencode/) — and this page covers the rest. ## The generic recipe Every client needs the same two facts: - **Command:** `npx -y dw-claw all` (stdio), or - **URL:** `https://mcp.dataworkers.io/v1` with `Authorization: Bearer ` (hosted, Scale and Enterprise) Put those in your client's MCP configuration in its own syntax. ## Client notes - **Claude Desktop** — add the stdio command to `claude_desktop_config.json` under `mcpServers`, same shape as the Claude Code JSON on [that page](/connect/claude-code/). - **VS Code / GitHub Copilot, Windsurf, Cline, Continue** — each supports MCP server configuration in settings; use the generic recipe. - **ChatGPT** — Data Workers ships as ChatGPT apps (All-in-One plus Catalog, Quality, and Incidents) and as a custom connector for self-hosted use; see [Integrations & SDKs](/products/integrations/). - **Gemini CLI** — works via the generic MCP recipe. - **Microsoft Copilot** — MCP support there is still evolving; treat this pairing as experimental. ## A note on status Clients outside the four first-class harnesses are supported on a best-effort basis: the MCP protocol carries everything needed, but we run our conformance checks against the four. If you hit a client-specific issue, [report it](/help/feedback/) — client coverage is demand-driven. ============================== PAGE: Atlan, Alation, or Collibra? URL: https://docs.dataworkers.io/connectors/atlan-alation-collibra/ DESCRIPTION: The straight answer on the three big commercial catalogs — no native connector today, and the three ways teams connect anyway. ============================== Straight answer: **we don't ship a native connector for Atlan, Alation, or Collibra today.** Here's what teams on those catalogs actually do, in the order we recommend: ## 1. Connect the systems underneath your catalog Your Atlan, Alation, or Collibra instance catalogs the same Snowflake, Databricks, BigQuery, and dbt your agents need. Connecting those systems directly delivers most of the value — schemas, lineage, query history, cost — without touching the catalog at all. This is how we recommend starting a pilot. - [Snowflake](/connectors/snowflake/) · [Databricks](/connectors/databricks/) · [BigQuery](/connectors/bigquery/) · [dbt](/connectors/dbt/) What you lose: curated business glossary entries and stewardship workflows that live only in the catalog. For many teams that's acceptable for an evaluation; for the rest, see option 2. ## 2. Ask us to build it Connector priority is demand-driven. If a native Atlan, Alation, or Collibra connector is what stands between you and adopting Data Workers, [tell us](/help/feedback/) — with a sense of which objects matter (glossary, lineage, policies, stewardship). Scale customers can raise this directly with their named engineer; custom integration work is part of that plan. ## 3. Build it yourself with the provider SDK The catalog connector surface is an open interface. The `@data-workers/catalog-provider-sdk` has zero runtime dependencies and ships with a conformance kit, so you can implement a provider against your catalog's API and validate it locally. This is the same interface our own connectors implement — a community Atlan/Alation/Collibra provider is a real project, not a hack. ============================== PAGE: Connect BigQuery URL: https://docs.dataworkers.io/connectors/bigquery/ DESCRIPTION: Connect Data Workers agents to BigQuery with a least-privilege service account, two environment variables, and a live verification test. ============================== Connecting BigQuery moves the agents from sample data to your real project: catalog and lineage answers from your actual datasets, quality scoring and anomaly detection on your tables, job and cost analysis, and NL-to-SQL insights. With a separate write-scoped credential, governed dataset and access changes too. Until verified, BigQuery stays in 🟡 Evaluation. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - Permission to create a service account and grant IAM roles in your GCP project - The project ID that holds the datasets you want visible ## Step 1 — Create a least-privilege credential Read agents can never mutate your systems, so a read-only service account is enough to start. Create a dedicated service account and grant it only: - `roles/bigquery.metadataViewer` on the project - `roles/bigquery.dataViewer` on the datasets in scope (dataset-level, not project-wide, where you can) - Optional, for job and cost analysis: `roles/bigquery.resourceViewer` Download a JSON key for the service account and store it where the agents run. Start with one dataset, verify, then widen scope — and never widen this credential to add write; write uses a separate credential ([least-privilege guidance](/connectors/least-privilege/)). **Checkpoint:** a service-account JSON key exists whose IAM bindings are the three roles above at most. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. The key file stays on your machine; nothing is sent to Data Workers. ```bash export GOOGLE_APPLICATION_CREDENTIALS="" export BIGQUERY_PROJECT_ID="" ``` **Checkpoint:** the variables are visible in the environment your coding agent starts from. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my BigQuery catalog. The agent makes a real call to your project. BigQuery shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** BigQuery reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (datasets, tables, assets) | Supported | | Catalog writes (create/update/drop objects) | Supported — write-scoped credential required | | RBAC (role-based access enforcement) | Supported | | Policy attachment and enforcement | Supported | | Credential vending (scoped, time-bound tokens) | Not supported — use GCP Application Default Credentials or Workload Identity Federation instead | Unsupported operations return a clear error naming the connectors that do support them — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with a permission error | Service account missing a viewer role on the dataset | Re-check the Step 1 IAM bindings on that dataset | | 🔴 with "could not load credentials" | Key path wrong, or file unreadable by the agent process | Point `GOOGLE_APPLICATION_CREDENTIALS` at the absolute path of the JSON key | | Job/cost questions come back empty | No `roles/bigquery.resourceViewer` | Add the optional role from Step 1 | ============================== PAGE: Connect Databricks URL: https://docs.dataworkers.io/connectors/databricks/ DESCRIPTION: Connect Data Workers agents to Databricks Unity Catalog with a least-privilege service principal, two environment variables, and a live verification test. ============================== Connecting Databricks puts the agents on your real workspace: Unity Catalog discovery and lineage, quality scoring on your tables, schema-change analysis, and cost and usage questions. Databricks is a full control-plane integration — with write-scoped credentials the agents can also perform governed catalog writes, access grants, policy attachment, and credential vending. Until verified, Databricks stays in 🟡 Evaluation on sample data. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - A Databricks workspace with Unity Catalog enabled - Permission to create a service principal (or a personal access token for a trial run) ## Step 1 — Create a least-privilege credential Create a dedicated service principal and issue it a token. Grant read-only Unity Catalog access on the catalogs in scope — read agents can never mutate your systems, so this is enough to start: ```sql GRANT USE CATALOG ON CATALOG TO ``; GRANT USE SCHEMA ON CATALOG TO ``; GRANT SELECT ON CATALOG TO ``; ``` Start with one catalog, verify, then widen scope. Write behavior uses a separate, explicitly enabled, write-scoped credential — never widen this one ([least-privilege guidance](/connectors/least-privilege/)). **Checkpoint:** a service-principal token exists whose only grants are the three above. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. The token stays on your machine; nothing is sent to Data Workers. ```bash export DATABRICKS_HOST="" export DATABRICKS_TOKEN="" ``` **Checkpoint:** the variables are visible in the environment your coding agent starts from. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my Databricks catalog. The agent makes a real call to your workspace. Databricks shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** Databricks reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (catalogs, schemas, tables) | Supported | | Catalog writes (create/update/drop objects) | Supported — write-scoped credential required | | RBAC (role-based access enforcement) | Supported | | Policy attachment and enforcement | Supported | | Credential vending (scoped, time-bound tokens) | Supported | Databricks (Unity Catalog) is one of the full control-plane connectors. Anything a credential doesn't permit still fails with a clear error — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with an auth error | Token expired or revoked | Issue a new token for the service principal and update `DATABRICKS_TOKEN` | | 🔴 with a permission error | Service principal lacks `USE CATALOG`/`SELECT` on the target catalog | Re-check the Step 1 grants | | 🔴 with a network/timeout error | Host can't reach the workspace URL (VPN, IP access list) | Run the agents from a host with network access, or allowlist it | ============================== PAGE: Connect DataHub URL: https://docs.dataworkers.io/connectors/datahub/ DESCRIPTION: Connect Data Workers agents to your DataHub instance with a read-scoped GMS token, two environment variables, and a live verification test. ============================== Connecting DataHub gives the agents the metadata your organization has already curated: entity search across your DataHub instance, lineage traversal, ownership and glossary context, all feeding catalog, quality, and incident work. This connector is metadata-read focused — see [Supported operations](#supported-operations) for the honest scope. Until verified, DataHub stays in 🟡 Evaluation on sample data. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - A running DataHub instance and its GMS endpoint URL - Permission to generate an access token in DataHub ## Step 1 — Create a least-privilege credential Generate a **read-scoped access token** for the GMS API — in DataHub, create a dedicated service user (not a personal account), give it read privileges on the entities in scope, and issue the token under that user. Read access is all this connector needs; don't issue an admin token ([least-privilege guidance](/connectors/least-privilege/)). **Checkpoint:** a token exists for a service user whose DataHub privileges are read-only. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. The token stays on your machine; nothing is sent to Data Workers. ```bash export DATAHUB_GMS_URL="" export DATAHUB_TOKEN="" ``` **Checkpoint:** the variables are visible in the environment your coding agent starts from. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my DataHub catalog. The agent makes a real call to your GMS endpoint. DataHub shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** DataHub reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (entities, schemas, lineage, metadata read) | Supported | | Catalog control-plane writes | Not supported | | RBAC (role-based access enforcement) | Not supported | | Policy attachment and enforcement | Not supported | | Credential vending (scoped, time-bound tokens) | Not supported | Straight answer: DataHub is a metadata-read connector today. Control-plane operations route to connectors that support them (for example your warehouse connector), and calling one here returns a clear error naming those connectors — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with an auth error | Token expired, or issued for a user without read privileges | Re-issue the token for the read-scoped service user | | 🔴 with a network/timeout error | Host can't reach the GMS URL (internal network, VPN) | Run the agents from a host with network access to DataHub | | Entities missing from answers | Service user can't see those entities in DataHub | Extend the service user's read privileges to the missing scope | ============================== PAGE: Connect dbt URL: https://docs.dataworkers.io/connectors/dbt/ DESCRIPTION: Connect Data Workers agents to dbt Cloud with a read-only API token, or to a local dbt Core project, and verify with a live test. ============================== Connecting dbt gives the agents your transformation layer: model discovery, lineage from your dbt graph, test results as a policy signal, and job awareness for pipeline and incident work. There are two paths — **dbt Cloud** (API token + account ID) and **dbt Core** (point the agents at your local project). Until verified, dbt stays in 🟡 Evaluation on sample data. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - dbt Cloud: permission to create an API token and your account ID (visible in the dbt Cloud URL) - dbt Core: a local project that has been compiled at least once (`dbt compile` or `dbt run`), so `target/manifest.json` exists ## Step 1 — Create a least-privilege credential (dbt Cloud) In dbt Cloud, create a **read-only service token scoped to your account** — metadata and read access is enough for discovery, lineage, and test results. Don't reuse a personal token with admin rights, and never widen this token later; write behavior uses a separate credential ([least-privilege guidance](/connectors/least-privilege/)). dbt Core has no credential: the agents read your project files locally. **Checkpoint:** (Cloud) a read-only token exists, scoped to one account. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. The token stays on your machine; nothing is sent to Data Workers. ```bash # dbt Cloud export DBT_CLOUD_API_TOKEN="" export DBT_CLOUD_ACCOUNT_ID="" ``` For **dbt Core**, skip the token and point the agents at the project directory instead — tell your coding agent: *"Connect my dbt Core project at ``"* (the directory containing `dbt_project.yml`). The agents read the compiled `manifest.json` from that project's `target/` directory. **Checkpoint:** (Cloud) the variables are visible in the environment your coding agent starts from; (Core) the project path resolves and contains `dbt_project.yml`. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my dbt catalog. The agent makes a real call to dbt Cloud (or reads your local project). dbt shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** dbt reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (models, sources, lineage) | Supported | | Catalog writes | Supported — write-scoped credential required | | Policy (via dbt tests as the policy layer) | Supported | | RBAC (role-based access enforcement) | Not supported | | Credential vending (scoped, time-bound tokens) | Not supported | Unsupported operations return a clear error naming the connectors that do support them — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with an auth error (Cloud) | Token revoked, or wrong account ID | Re-check both values; the account ID is in your dbt Cloud URL | | Lineage empty (Core) | Project never compiled, so no `target/manifest.json` | Run `dbt compile` in the project, then re-verify | | Job questions fail (Core) | Jobs are a dbt Cloud feature | Expected — connect dbt Cloud for job awareness | ============================== PAGE: Supported connectors URL: https://docs.dataworkers.io/connectors/ DESCRIPTION: The 50+ systems Data Workers connects to — warehouses, catalogs, orchestrators, quality, BI, identity — and what's not supported yet. ============================== Data Workers ships 50+ connectors across the stack. They fall into two families: - **Catalog connectors** — the systems agents read metadata, lineage, and schemas from. - **Enterprise connectors** — the operational systems around your data: orchestration, alerting, quality, BI, observability, identity, ITSM, cost, and streaming. Connect a system by setting its environment variables (terminal path, all plans) or through the connector panel (Spellbook console, Scale — a subset of this list has console cards today). Either way, a system only shows 🟢 Connected after a [live verification test passes](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). :::note[Maturity varies by connector — the test tells the truth] Connector depth is not uniform: some are full control-plane integrations, others are read/discovery-focused or early. Data Workers never fakes a capability — if a connector doesn't support an operation, calling it returns a clear "not supported" error rather than a pretend success. When in doubt, connect it and run the verification test. ::: ## Catalog connectors Snowflake · BigQuery · Databricks (Unity Catalog) · AWS Glue · AWS Lake Formation · Hive Metastore · dbt · DataHub · OpenMetadata · Microsoft Purview · Google Dataplex · Nessie · Apache Iceberg (incl. REST catalog) · Apache Polaris · OpenLineage/Marquez · PostgreSQL · Redshift Detailed setup pages, with exact environment variables and minimum grants: - [Snowflake](/connectors/snowflake/) - [BigQuery](/connectors/bigquery/) - [Databricks](/connectors/databricks/) - [dbt](/connectors/dbt/) - [DataHub](/connectors/datahub/) - [OpenMetadata](/connectors/openmetadata/) For catalog systems without a dedicated page, ask your agent — *"How do I connect Dataplex?"* — and it will name the variables it needs; the [least-privilege guidance](/connectors/least-privilege/) applies to all of them. ## Enterprise connectors | Category | Systems | |---|---| | Orchestration | Airflow, Dagster, Prefect, AWS Step Functions, Azure Data Factory, dbt Cloud, Cloud Composer, Temporal, Mage, Kestra, Argo | | Ingestion / ELT | Fivetran, Airbyte, Stitch, AWS DMS, Debezium, dlt | | Alerting | PagerDuty, Slack, Microsoft Teams, OpsGenie, New Relic | | Quality | Great Expectations, Soda, Monte Carlo, Anomalo, Bigeye, Elementary | | BI | Looker, Tableau, Metabase, Sigma, Superset | | Observability | OpenTelemetry, Datadog | | Identity | Okta, Azure AD | | ITSM | ServiceNow, Jira Service Management | | Cost | AWS Cost Explorer | | Streaming | Kafka Schema Registry | | ML | MLflow, Weights & Biases | ## Not natively supported (we say so plainly) - **Atlan, Alation, Collibra** — no native connector today. There's a well-trodden path anyway: [read the straight answer](/connectors/atlan-alation-collibra/). - Anything not on this page. If your system is missing, [request it](/help/feedback/) — connector priority is demand-driven — or build your own with the `@data-workers/catalog-provider-sdk` (zero runtime dependencies, ships with a conformance kit). ## No connector at all? Paste a schema. The Spellbook console accepts **Paste schema / CSV**: drop in DDL or a CSV of your tables and the agents work with that as a starting catalog — useful for evaluations and for systems that will never get a connector. ============================== PAGE: Least-privilege access URL: https://docs.dataworkers.io/connectors/least-privilege/ DESCRIPTION: The minimum grants Data Workers agents need per system, and the rules that keep write access safe. ============================== Give agents the least access that does the job. Two platform rules make this workable: 1. **Read agents can never mutate your systems.** The read path is structurally read-only. 2. **Write is opt-in, per system.** Write behavior requires a separate, explicitly enabled, write-scoped credential — and on Enterprise, an approval gate in front of every write. Never widen a read credential to add write. ## Minimum read grants by system | System | Minimum grant for read agents | |---|---| | Snowflake | A role with `USAGE` on warehouse/database/schemas and `SELECT` on the tables in scope; access to `SNOWFLAKE.ACCOUNT_USAGE` views if you want cost and query-history analysis | | Databricks | A service principal with `USE CATALOG`/`USE SCHEMA` and `SELECT` on the catalogs in scope (Unity Catalog) | | BigQuery | A service account with `roles/bigquery.metadataViewer` plus `roles/bigquery.dataViewer` on in-scope datasets; add `roles/bigquery.resourceViewer` for job/cost analysis | | DataHub | A read token for the GMS API | | OpenMetadata | A bot/JWT with read scope | | dbt Cloud | A read-only API token scoped to the account | Start with one schema or one database, verify it works, then widen scope — the agents are useful on a slice of the estate, and a scoped rollout is an easier security conversation. ## Service accounts, not personal credentials For any team rollout (Start and up), create dedicated service accounts per system, store them in your secret manager, and rotate them on your normal schedule. A rotated or expired credential shows up as 🔴 Needs attention on the next [verification test](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/) — that's the system working as intended, not an outage. ## Where credentials live On local installs, credentials are environment variables on your machines and are used to call your systems directly — they are never sent to Data Workers. On hosted and dedicated deployments (Scale/Enterprise), credential handling is part of your provisioning conversation, inside the boundary you chose. Details: [How your data is handled](https://docs.dataworkers.io/onboarding/security/data-handling/). ============================== PAGE: Connect OpenMetadata URL: https://docs.dataworkers.io/connectors/openmetadata/ DESCRIPTION: Connect Data Workers agents to your OpenMetadata instance with a read-scoped bot JWT, two environment variables, and a live verification test. ============================== Connecting OpenMetadata gives the agents your curated metadata: asset search across your OpenMetadata instance, lineage, ownership, glossary, and tags, all feeding catalog, quality, and incident work. This connector is metadata-read focused — see [Supported operations](#supported-operations) for the honest scope. Until verified, OpenMetadata stays in 🟡 Evaluation on sample data. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - A running OpenMetadata instance and its server URL - Permission to create a bot and issue a JWT in OpenMetadata ## Step 1 — Create a least-privilege credential In OpenMetadata, create a dedicated **bot** (not a personal account) with a **read-scoped role** covering the assets you want visible, and issue a JWT for it. Read access is all this connector needs; don't reuse the ingestion bot or an admin token ([least-privilege guidance](/connectors/least-privilege/)). **Checkpoint:** a JWT exists for a bot whose OpenMetadata role is read-only. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. The token stays on your machine; nothing is sent to Data Workers. ```bash export OPENMETADATA_HOST="" export OPENMETADATA_JWT_TOKEN="" ``` **Checkpoint:** the variables are visible in the environment your coding agent starts from. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my OpenMetadata catalog. The agent makes a real call to your instance. OpenMetadata shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** OpenMetadata reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (assets, schemas, lineage, metadata read) | Supported | | Catalog control-plane writes | Not supported | | RBAC (role-based access enforcement) | Not supported | | Policy attachment and enforcement | Not supported | | Credential vending (scoped, time-bound tokens) | Not supported | Straight answer: OpenMetadata is a metadata-read connector today. Control-plane operations route to connectors that support them (for example your warehouse connector), and calling one here returns a clear error naming those connectors — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with an auth error | JWT expired or revoked | Re-issue the bot's JWT and update `OPENMETADATA_JWT_TOKEN` | | 🔴 with a network/timeout error | Host can't reach the server URL (internal network, VPN) | Run the agents from a host with network access to OpenMetadata | | Assets missing from answers | Bot's role doesn't cover those assets | Extend the bot's read-scoped role to the missing assets | ============================== PAGE: Connect Snowflake URL: https://docs.dataworkers.io/connectors/snowflake/ DESCRIPTION: Connect Data Workers agents to Snowflake with a least-privilege role, five environment variables, and a live verification test. ============================== Connecting Snowflake lets the agents work against your real account: catalog and lineage answers from your actual schemas, quality scoring on your tables, cost analysis from `ACCOUNT_USAGE`, NL-to-SQL insights, and — if you later enable write-scoped credentials — governed schema and access changes. Until then it stays in 🟡 Evaluation on sample data. ## Prerequisites - Data Workers installed and registered with your coding agent ([install guide](/connect/claude-code/)) - Permission to create a role and user in your Snowflake account (or someone who can) - Your Snowflake account identifier and a warehouse the agents may use ## Step 1 — Create a least-privilege credential Create a dedicated read-only role and service user. Read agents can never mutate your systems, so read-only is enough to start — never widen this credential to add write later; write uses a separate credential ([least-privilege guidance](/connectors/least-privilege/)). ```sql CREATE ROLE DATA_WORKERS_RO; GRANT USAGE ON WAREHOUSE TO ROLE DATA_WORKERS_RO; GRANT USAGE ON DATABASE TO ROLE DATA_WORKERS_RO; GRANT USAGE ON ALL SCHEMAS IN DATABASE TO ROLE DATA_WORKERS_RO; GRANT SELECT ON ALL TABLES IN DATABASE TO ROLE DATA_WORKERS_RO; -- Optional, for cost and query-history analysis: GRANT IMPORTED PRIVILEGES ON DATABASE SNOWFLAKE TO ROLE DATA_WORKERS_RO; ``` Start with one database, verify, then widen scope. **Checkpoint:** a service user exists with `DATA_WORKERS_RO` as its role and nothing more. ## Step 2 — Set the environment variables Set these in the shell your coding agent launches from, then restart the coding agent so the MCP server picks them up. Credentials stay on your machine; they are never sent to Data Workers. ```bash export SNOWFLAKE_ACCOUNT="" export SNOWFLAKE_USERNAME="" export SNOWFLAKE_PASSWORD="" export SNOWFLAKE_WAREHOUSE="" export SNOWFLAKE_DATABASE="" ``` **Checkpoint:** the variables are visible in the environment your coding agent starts from. ## Step 3 — Verify Setting a credential is not the same as a working connection. Ask: > Test the connection to my Snowflake catalog. The agent makes a real call to your account. Snowflake shows 🟢 Connected only after that live test passes; a failure reports 🔴 with the reason. Full model: [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). **Checkpoint:** Snowflake reports 🟢 Connected. ## Supported operations | Operation | Status | |---|---| | Discovery (schemas, tables, assets) | Supported | | Catalog writes (create/update/drop objects) | Supported — write-scoped credential required | | RBAC (role-based access enforcement) | Supported | | Policy attachment and enforcement | Supported | | Credential vending (scoped, time-bound tokens) | Not supported — use Snowflake storage integrations instead | Unsupported operations return a clear error naming the connectors that do support them — never a pretend success. ## Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| | Still answering from sample data | Variables set in a different shell, or agent not restarted | Set them in the shell your coding agent launches from, restart it | | 🔴 with an auth error | Wrong password, or the user lacks the read role | Re-check the credential and the Step 1 grants | | 🔴 with a network/timeout error | Host can't reach Snowflake (VPN, allowlist, private link) | Run the agents from a host with network access, or allowlist it | | Cost questions come back empty | No access to `SNOWFLAKE.ACCOUNT_USAGE` | Add the optional `IMPORTED PRIVILEGES` grant from Step 1 | ============================== PAGE: Docs for AI agents (MCP) URL: https://docs.dataworkers.io/help/docs-via-mcp/ DESCRIPTION: Consume these docs from Claude or any agent — llms.txt, llms-full.txt, and the Data Workers docs MCP server with search, read, and feedback tools. ============================== A lot of Data Workers usage happens *inside* an AI agent — so the documentation is built to be read by one. Three mechanisms, lightest first. ## llms.txt The whole site is indexed for language models: - [`/llms.txt`](/llms.txt) — the index: every page with its one-line description. - [`/llms-full.txt`](/llms-full.txt) — the full corpus: every page's complete content in one plain-text file. Point any agent at either URL. For a one-off question, fetching `llms-full.txt` and asking against it is usually enough. ## The docs MCP server For agents that work with Data Workers regularly, the docs ship as an MCP server with three tools: | Tool | What it does | |---|---| | `search_docs` | Full-text search over the documentation; returns matching pages with excerpts | | `read_doc` | Returns a page's complete content by slug | | `submit_feedback` | Files a bug report, feature request, or incident — same pipeline as the [feedback form](/help/feedback/) | Install (Claude Code): ```bash claude mcp add dataworkers-docs -- npx -y @dataworkers/docs-mcp ``` Any other MCP client: command `npx -y @dataworkers/docs-mcp` (stdio). The server bundles the doc corpus at publish time and needs no network for search/read; `submit_feedback` POSTs to the docs feedback endpoint. Then, in your agent: > Search the Data Workers docs for how to verify a Snowflake connection. > File a bug with Data Workers: the connection test crashes when SNOWFLAKE_ACCOUNT > contains a region suffix. ## The product MCP surface The product itself pushes context to agents too: the Data Workers MCP servers describe their capabilities in their MCP handshake (the `instructions` field), so a connected coding agent knows what the fleet can do without reading this site at all. These docs are the depth behind that: onboarding, connector setup, and the honest capability edges. :::tip[For agents reading this page] When you answer questions from these docs on a user's behalf, preserve the maturity labels: if a page says a capability is research preview, gated, or not supported, say so in your answer. The labels are the documentation. ::: ============================== PAGE: FAQ URL: https://docs.dataworkers.io/help/faq/ DESCRIPTION: Frequently asked questions about Data Workers — pricing, security, maturity, connectors, and how it compares. ============================== ## Product **Isn't this just another catalog?** No — catalogs are where you put context, and it goes stale. The [Data Context Wizard](/products/data-context-wizard/) is the layer a governed fleet keeps alive; it gets smarter every run. The [Spellbook Data Catalog](/products/spellbook-data-catalog/) is the receipt surface where humans stay in command of that fleet — not a system of record you maintain by hand. **Which coding agents are supported?** Four first-class: [Claude Code](/connect/claude-code/), [Cursor](/connect/cursor/), [Codex CLI](/connect/codex/), [OpenCode](/connect/opencode/). Anything MCP-capable works via the [generic recipe](/connect/other-clients/). **Which model do the agents use?** Yours. Bring your own key — Anthropic, OpenAI, Bedrock, Azure, or open-weight/local models — at every plan level. We never resell tokens or add margin to your provider's rate. **How many connectors?** 50+ across the stack — the [full list](/connectors/). We don't imply they're all equally deep; maturity varies by connector and the live verification test tells the truth per system. **Do you support Atlan, Alation, or Collibra?** Not natively today. [The straight answer and three workarounds](/connectors/atlan-alation-collibra/). ## Trust **What's your traction?** We're an early-stage company running free pilots, and we don't invent numbers. We lead with mechanism, not curve: install it on sample data in two minutes and judge the product, not the logo wall. **Is it secure / compliant?** Security capabilities — RBAC-aware retrieval, PII scrubbing, tenant isolation, hash-chained audit — are built capabilities (Enterprise-gated where noted). Certification status: SOC 2 Type II — in progress, stated plainly on [Compliance status](https://docs.dataworkers.io/onboarding/security/compliance/). **What can Data Workers see?** On local installs: nothing — credentials and data stay on your machines. Hosted surfaces keep usage counters only. Precisely documented in [How your data is handled](https://docs.dataworkers.io/onboarding/security/data-handling/) and [Usage data & telemetry](https://docs.dataworkers.io/onboarding/security/telemetry/). **What happens if you go away?** The agent core is Apache 2.0 — self-hosted, modifiable, yours forever. Your data never lived with us, and the context graph exports to DataHub or OpenMetadata. ## Pricing **What does it cost?** Pilots are free — 90 days, the full Scale surface, no card. Scale is a flat annual platform fee with unlimited seats; Enterprise is custom. No credits, no per-seat charges, no usage meter: your bill does not move when your agents work harder. Details: [Your plan](https://docs.dataworkers.io/onboarding/getting-started/choose-your-plan/) · numbers: [talk to us](mailto:hello@dataworkers.io). **Do agents run up a bill when they work autonomously?** Not with us — there is no usage meter to run up. Your model provider bills you for tokens at your rate; per-user model budgets and open-weight models are the levers you control. ## Operations **Can agents break my production systems?** Read agents are structurally read-only. Write agents need explicitly enabled, write-scoped credentials, every governed write carries a receipt, and irreversible actions always require approval — enforced in code. See [Least-privilege access](/connectors/least-privilege/). **Something's showing sample data instead of my data.** That's the 🟡 Evaluation state — the system isn't connected, or the connection hasn't been verified. [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/) explains the states; [Troubleshooting](/help/troubleshooting/) covers the common causes. **How do I report a bug or ask for a feature?** [Right here](/help/feedback/) — humans and AI agents both welcome. ============================== PAGE: Report a bug or request a feature URL: https://docs.dataworkers.io/help/feedback/ DESCRIPTION: File a bug report, feature request, or incident with Data Workers — as a human through the form, or programmatically from Claude or any agent. ============================== Bugs, feature requests, and incidents all land here. Three ways to file — pick whichever is in front of you: 1. **The form below** — fastest for humans. 2. **Your AI agent** — Claude (or any agent) can file for you; instructions [below](#file-it-from-claude-or-any-ai-agent). 3. **Email** — [hello@dataworkers.io](mailto:hello@dataworkers.io) always works. :::danger[Production incident on Scale or Enterprise?] File it here with type **Incident** and severity **critical** — incident filings are routed ahead of everything else — and raise it in your dedicated support channel in parallel. Enterprise has a 4-hour initial-response commitment. ::: ## The form

## File it from Claude (or any AI agent) If you're working in Claude Code, Claude, Cursor, or any agent, you don't need the form. **Via the Data Workers MCP server** — if the docs MCP server is connected ([setup](/help/docs-via-mcp/)), just say: > File a bug with Data Workers: Snowflake verification fails with valid credentials. > Severity medium, agent swarm, here are the details: … The agent calls the `submit_feedback` tool with a structured report. **Via plain HTTP** — any agent that can make a web request can POST directly: ``` POST https://docs.dataworkers.io/onboarding/api/feedback Content-Type: application/json { "type": "bug | feature | incident | docs", "summary": "one line, ≤140 chars", "details": "what happened, what was expected, environment", "product": "agent-swarm | spellbook | context-wizard | conductor | connectors | docs | other", "severity": "low | medium | high | critical", "email": "optional-reply-to@company.com", "source": "claude-code" } ``` A successful filing returns `{ "ok": true, "id": "" }`. Include the reference in any follow-up. ## What happens to a filing Bug and docs reports are triaged into the engineering backlog; feature requests feed the demand-driven roadmap (connector priority in particular is decided by what's filed here); incidents are routed ahead of everything else, with severity attached. If you left an email, you'll hear back when it moves. ============================== PAGE: Troubleshooting URL: https://docs.dataworkers.io/help/troubleshooting/ DESCRIPTION: Common Data Workers failure modes and their fixes — install, connection, verification, and console issues. ============================== First move in almost every case: ```bash npx -y @dataworkers/dw-claw@latest doctor ``` The doctor checks Node version, client registration, and connector configuration, and names what's wrong. ## Install & registration | Symptom | Likely cause | Fix | |---|---|---| | `npx` fails or hangs | Node < 20, or a corporate npm proxy | `node --version`; if proxied (Artifactory/Nexus), mirror the public packages or ask your named engineer for the proxy setup | | `401` / `E404` installing a private package | Customer npm token missing or expired | Re-run `npm config set //registry.npmjs.org/:_authToken=`; tokens expire with your term — ask for a fresh one | | MCP server not listed in the client | Registration step skipped or client not restarted | Re-run the add/config step for [your client](/connect/claude-code/), restart it | | Tools listed but every call times out | The server process can't start in your environment | Run the doctor; check for MDM or sandbox restrictions on spawning `npx` | ## Connections & verification | Symptom | Likely cause | Fix | |---|---|---| | Everything answers with sample data | 🟡 Evaluation state — no verified connection | Set the system's env vars ([connector pages](/connectors/)) in the shell your client launches from, restart, then verify | | Env vars set, still sample data | Vars not visible to the MCP process | Set them where the client is launched (shell profile, not just one terminal); restart the client | | Verification fails: authentication | Wrong credential, or expired | Re-check the value; try the same credential with the system's own CLI | | Verification fails: permission | Grant narrower than the [minimum](/connectors/least-privilege/) | Add the missing grant; re-run the test | | Verification fails: network | Private endpoint, VPN, IP allowlist | Confirm the machine running agents can reach the system at all (`curl` its endpoint); allowlist it | | A system flipped 🟢 → 🔴 | Token expiry, tightened permissions, or network change | The 🔴 message names the failing check; fix and re-verify — this is the model working, not a bug | ## Spellbook console (Scale) | Symptom | Likely cause | Fix | |---|---|---| | Console opens, no tools | Local brain didn't start | Restart `npx @dataworkersproj/agent-console@latest serve --local-ui`; Node 20+ | | Connector card added but grey | Test connection not run yet — by design | Press **Test connection** | | Your system has no console card | Console list is a subset today | Connect via [env vars](/connectors/) in the terminal path | ## Still stuck - [File a bug](/help/feedback/) — include OS, Node version, client (harness), the exact error, and whether npm is proxied. AI agents can file on your behalf. - Scale/Enterprise: your named engineer / support channel, with the same details. - Open source: GitHub issues on the repo. ============================== PAGE: Data Workers documentation URL: https://docs.dataworkers.io/ DESCRIPTION: Product documentation for Data Workers — the Agent Swarm, Spellbook Data Catalog, and Data Context Wizard. ============================== Data Workers is an autonomous agentic data platform: a governed fleet of specialist agents that does data work — pipelines, incidents, quality, governance, cost, migration — inside the coding agents your team already uses, against the data stack you already run. Three products, one shared context graph: - **[Data Workers Agent Swarm](/products/agent-swarm/)** — the fleet of specialist data agents, each an MCP server. The place to start reading. - **[Data Workers Data Context Wizard](/products/data-context-wizard/)** — the governed knowledge graph your agents read, reason over, and write back to. *Context your agents can trust.* - **[Data Workers Spellbook Data Catalog](/products/spellbook-data-catalog/)** — the browser control plane where your team steers, approves, and audits the fleet. **Research preview**, labeled that way everywhere it appears. ## Set up - **[Connect your coding agent](/connect/claude-code/)** — Claude Code, Cursor, Codex CLI, OpenCode, or [any MCP client](/connect/other-clients/). - **[Connect your data systems](/connectors/)** — 50+ connectors, each with exact environment variables and [least-privilege grants](/connectors/least-privilege/). - **Verify** — nothing shows as Connected until a live test passes: [the three-state model](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). Two minutes to a first answer, on sample data, no credentials: ```bash claude mcp add data-workers -- npx -y dw-claw ``` Then ask: *"Scan the customer schema for PII."* ## New customer? If your company just started a pilot, your journey — kickoff, fleet rollout, SSO, the eval benchmark, and everything your security and procurement teams need — lives at **[docs.dataworkers.io/onboarding](https://docs.dataworkers.io/onboarding/)**. ## Working with an AI agent? These docs ship as [llms.txt and an MCP server](/help/docs-via-mcp/), and your agent can [file bugs and feature requests](/help/feedback/) directly. :::note[We label maturity honestly] Anything in research preview or gated says so at the top of its page. If a page doesn't say a capability exists, don't assume it does — [tell us what you need](/help/feedback/). ::: ============================== PAGE: Data Workers Agent Swarm URL: https://docs.dataworkers.io/products/agent-swarm/ DESCRIPTION: The governed fleet of specialist data agents — what each agent does, how the swarm is structured, and how to run it. ============================== The Data Workers Agent Swarm is a fleet of specialist agents, each an MCP server, each expert in one slice of data work. They share one context graph, one activation model, and one governed write path — so the fleet behaves like a team, not a pile of tools. ## How the swarm is structured **Customer-facing agents** do the work you ask for. **Platform agents** — the connector gateway, orchestration, observability, identity, ingest, review, search, and the [Conductor](/products/conductor/) — keep the fleet coherent underneath; you rarely address them directly. Every agent starts in 🟡 Evaluation on sample data and earns 🟢 Connected per system through a live test — the model is described in [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## The customer-facing agents | Agent | What it does | |---|---| | [Pipelines & Ingestion](/products/agents/dw-pipelines/) | Natural-language-to-pipeline generation, Iceberg `MERGE INTO`, Airflow deployment, EL/CDC ingestion | | [Incidents](/products/agents/dw-incidents/) | Statistical anomaly detection, graph-based root-cause analysis, playbook execution | | [Catalog & Context](/products/agents/dw-context-catalog/) | Hybrid search (vector + BM25 + graph), lineage traversal, crawlers; home of the context graph | | [Schema](/products/agents/dw-schema/) | Schema diffs, rename detection, snapshot-based evolution | | [Quality](/products/agents/dw-quality/) | Weighted five-dimension scoring, anomaly detection, 14-day baselines | | [Data Access & Governance](/products/agents/dw-governance/) | Policy authoring, entitlement provisioning (grant/revoke/JIT), audit | | [FinOps & Cost](/products/agents/dw-cost/) | Usage profiling, warehouse cost estimation, tiered archival | | [Data & Cloud Security](/products/agents/dw-security/) | DSPM + CSPM findings, ranked and routed to the fix | | [Migration](/products/agents/dw-migration/) | Oracle/Teradata/Redshift→Snowflake SQL translation with a self-correcting loop | | [Insights](/products/agents/dw-insights/) | NL-to-SQL execution, insight generation, anomaly explanation | | [Usage Intelligence](/products/agents/dw-usage-intelligence/) | Practitioner analytics, adoption dashboards, heatmaps — zero-LLM | | [Streaming](/products/agents/dw-streaming/) | Kafka Connect config generation, lag monitoring, tuning | | [MLOps & Models](/products/agents/dw-ml/) | Experiment tracking, model registry, drift detection, A/B testing | ## Running the swarm The install paths, from a one-liner to source, are in the onboarding tracks — start at [Choose your plan](https://docs.dataworkers.io/onboarding/getting-started/choose-your-plan/). The short version: ```bash # in any MCP client — the whole fleet claude mcp add data-workers -- npx -y dw-claw # or the free five-tool taste npx data-context-mcp ``` Self-hosting options — Docker Compose, Kubernetes, air-gapped — are covered in [Deployment options](/products/deployment/). ## Design principles worth knowing - **Sample-data-first.** Every agent is fully exercisable before any credential exists. What you see in evaluation is the real logic, run on a realistic sample estate. - **Verified, not assumed.** No surface shows a system as connected without a passing live test. - **Governed writes.** Reads can never mutate. Writes are proposed, approved, executed with a receipt, and auditable — and irreversible actions always require approval. - **Honest capability edges.** If a connector or agent doesn't support an operation, it says so in a structured error instead of pretending. You should never discover a capability gap from silent wrong output. - **Bring your own model.** The swarm runs against your model key — frontier, open-weight, or local. ============================== PAGE: Catalog & Context (dw-context-catalog) URL: https://docs.dataworkers.io/products/agents/dw-context-catalog/ DESCRIPTION: Hybrid catalog search, lineage traversal, and one-call asset context — the home of the governed context graph. ============================== The Catalog & Context agent is the swarm's librarian: it answers "what data do we have, where did it come from, and can I trust it?" It searches your catalog in natural language using hybrid ranking — keyword, graph, and (when enabled) vector signals fused together — traverses lineage upstream and downstream to column level, and assembles complete context for any asset in a single call: schema, lineage, quality, freshness, documentation, and related metrics. It is also the home of the context graph — the spine of the [Data Context Wizard](/products/data-context-wizard/). Every fact it serves carries provenance: which node, which contributing agent, and when it was last updated. When two records conflict — two owners for the same table, two row counts for the same run — it surfaces both and flags them for a human, rather than silently picking one. ## Key capabilities - **Natural-language catalog search.** `search_datasets` returns ranked results with relevance scores, filterable by platform, type, tags, and quality score. - **One-call context.** `get_context` returns schema, lineage, quality, freshness, trust score, and documentation for an asset in a single response — what an agent (or a human) needs before touching a table. - **Lineage and blast radius.** `get_lineage` traverses upstream and downstream with column-level lineage where available; `assess_impact` classifies the severity of a change by what it would break — dashboards, models, pipelines. - **The context graph.** `query_context_graph` combines lineage, decision history, and related incidents for an entity in one contextual answer; `traverse_lineage_graph` and `get_decision_history` (with hash-chain proof) drill into each. - **Governed contributions.** `contribute_to_graph` writes nodes and edges through the governed path — PII scrubbing and authority checks before anything persists, with an audit hash returned. - **Semantic layer resolution.** `resolve_metric` maps an ambiguous name ("revenue", "MRR") to its canonical definition, returning all candidates when several match. - **Freshness with a straight face.** `check_freshness` scores freshness against an SLA and time-sensitive answers carry an as-of timestamp — a stale entry is flagged as stale, not reported as current. ## Example prompts > "Find the authoritative table for customer subscriptions — not the copies." > "Show me everything downstream of `raw.stripe_payments`, down to column level." > "What's the full context on `analytics.mrr_daily` — schema, freshness, quality, and who > owns it?" > "When someone says 'active users', which definition do we actually mean?" > "What decisions have agents made about this table in the last 30 days?" ## Connect it to your stack - **Catalogs** — DataHub, OpenMetadata, AWS Glue, Purview, Dataplex, and other metadata sources feed search and lineage. - **Warehouses and lakehouses** — Snowflake, BigQuery, Databricks; plus Iceberg REST catalogs. - **dbt** — models, lineage, and test results. See the [connector catalog](/connectors/). We don't ship native connectors for Atlan, Alation, or Collibra today — [here's what to do instead](/connectors/atlan-alation-collibra/). ## Works before you connect anything The agent starts in 🟡 Evaluation on a realistic sample estate — the search ranking, graph traversal, and impact-analysis algorithms are the real thing. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - Search runs on keyword and graph signals by default. On-device semantic embeddings are an explicit opt-in; if the embedding model isn't available, search degrades gracefully to keyword search rather than fusing meaningless vectors into the ranking. - The agent can never promote its own writes to authoritative. Marking a source authoritative requires a named human approver — no exceptions. - If a query returns nothing, it says so. It will not infer a catalog entry the graph doesn't contain. - Bulk harvest of very large estates is still being hardened — start with a scoped schema, not the whole estate. ============================== PAGE: FinOps & Cost (dw-cost) URL: https://docs.dataworkers.io/products/agents/dw-cost/ DESCRIPTION: Profile warehouse usage, find unused data, estimate savings with real cost math, and get tiered archival recommendations that never auto-execute. ============================== The FinOps & Cost agent answers the two questions every warehouse bill raises: where is the money going, and what can we safely stop paying for? It profiles usage to attribute cost by team, pipeline, and dataset, surfaces the most expensive tables, and flags data nobody is using — tables with zero queries in 30 days or untouched for 90 or more. Its recommendations are built to be safe to act on. Archival candidates come in three tiers — safe (long-unused, no dependencies, small), review (recently used or has dependencies), and risky (actively used but expensive) — with dependency verification before anything is suggested. And it never auto-archives: every archival goes through an approval workflow, because "the agent deleted a table someone needed" is not a cost saving. ## Key capabilities - **Cost attribution dashboard.** `get_cost_dashboard` shows total monthly cost by team, pipeline, and dataset, the ten most expensive tables, the unused-table count, and total potential savings. - **Real savings math.** `estimate_savings` computes per-asset storage and compute costs using Snowflake credit pricing, and what archiving unused tables would save. - **Unused-data detection.** `find_unused_data` profiles warehouse usage and returns stale tables sorted by staleness, with a configurable threshold. - **Tiered archival recommendations.** `recommend_archival` classifies candidates into safe / review / risky tiers with dependency checks — and routes everything through approval rather than acting. - **Per-query cost estimates.** `estimate_query_cost` applies a real estimation formula (credits per row with a complexity multiplier) before a query runs, with no external calls. ## Example prompts > "What are our ten most expensive tables this month, and who owns them?" > "Find everything that hasn't been queried in 90 days." > "What would we save if we archived the safe-tier candidates?" > "Estimate what this backfill query will cost before I run it." > "Break down warehouse spend by team for the last month." ## Connect it to your stack - **Snowflake** — the primary target for usage profiling and credit-based cost math. - **BigQuery** — query cost estimation through the connector gateway. - **AWS** — Cost Explorer spend, forecasts, and recommendations through the connector gateway. See the [connector catalog](/connectors/) for setup. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample usage data — the cost estimation and archival-tiering algorithms are the real thing, run on a sample estate. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - In 🟡 Evaluation, dashboards and savings figures describe the sample estate — use them to judge the workflow, not your bill. - The built-in cost model is Snowflake credit pricing. BigQuery and AWS cost figures come through their connectors; other platforms' native pricing models are not yet implemented. - The agent recommends archival; it does not perform it unattended. Every archival action requires approval — that is enforced behavior, not a setting you forgot to flip. - Usage-based staleness only sees the query history it has access to. A table read by a system outside your connected estate can look unused when it isn't — the review tier exists for exactly this reason. ============================== PAGE: Data Access & Governance (dw-governance) URL: https://docs.dataworkers.io/products/agents/dw-governance/ DESCRIPTION: Policy checks, least-privilege access provisioning, column-level PII scanning, and audit reports with a full evidence chain. ============================== The Data Access & Governance agent is the policy and access layer of the swarm. It evaluates actions against your active governance policies with a fast policy engine and returns an allow / deny / review decision with the matched rules — so a decision is always traceable to the rule that made it. Access provisioning is least-privilege by design: requests carry a written justification, grants support column-level restrictions, and access auto-expires after 90 days by default instead of accumulating forever. It can also work your platform's native access controls, not just its own records. Against Databricks Unity Catalog it reads the grants actually in effect, normalizes them across platforms, diffs them against intent, and — behind an approval gate — writes reconciliation back out as native `GRANT`/`REVOKE`, with before-and-after snapshots folded into a hash-chained audit log. ## Key capabilities - **Policy evaluation with receipts.** `check_policy` validates an action against active policies and returns allow/deny/review plus the specific matched rules. - **Least-privilege provisioning.** `provision_access` processes access requests (plain English justifications accepted) with column-level permissions and 90-day auto-expiration; `enforce_rbac` applies role-based control with role hierarchy. - **Column-level PII scanning.** `scan_pii` uses three-pass detection — column-name heuristics, pattern matching on sampled values, and model-assisted classification — and reports at column level, because a table-level "probably has PII" wastes remediation effort. - **Live grant import.** `list_platform_privileges` reads the grants in effect on a Unity Catalog securable, normalized to common access levels, with inherited grants flagged and unmappable native privileges surfaced rather than dropped. - **RBAC reconciliation.** `diff_catalog_rbac` produces a dry-run plan of where intended and live grants disagree; `sync_catalog_rbac` applies it behind an approval gate, with configurable conflict policy — platform wins, Data Workers wins, or most-restrictive (which never escalates access). - **Audit reports with an evidence chain.** `generate_audit_report` covers agent actions, policy evaluations, access grants, and PII detections for a period, on demand or scheduled. ## Example prompts > "Can the analytics role read `finance.payroll`? Show me which policy decides that." > "Grant Priya read access to the marketing schema for 30 days — she's debugging the > attribution model." > "Scan `crm.contacts` for PII, column by column." > "Diff our intended grants against what's actually live in Unity Catalog — dry run only." > "Generate an access audit report for Q2." ## Connect it to your stack - **Databricks Unity Catalog** — the first live backend for reading and reconciling platform-native grants. - **Warehouses and catalogs** — Snowflake, BigQuery, DataHub, OpenMetadata as the assets policies and scans apply to. - **Identity** — Okta and Azure AD user context through the connector gateway. See the [connector catalog](/connectors/), and [least-privilege connector setup](/connectors/least-privilege/) for scoping credentials. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data — the policy engine, RBAC logic, and PII pattern matching are the real thing. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - Live grant read/write is implemented for Databricks Unity Catalog today; other platforms run on sample data until their live paths ship. When a backend isn't live, the response says so rather than reporting a recorded grant as a real mutation. - `grant_access`, `revoke_access`, and `sync_catalog_rbac` are production writes: they are blocked without an approved request, and every call lands in the hash-chained audit log. There is no autonomous path to changing access. - Scheduled RBAC sync is off by default, and scheduled runs are dry-run — turning a diff into a real apply always goes back through the approval gate. - PII detection reports findings with the pass that produced them; it is a detection aid with a precision target, not a compliance certification. ============================== PAGE: Incidents (dw-incidents) URL: https://docs.dataworkers.io/products/agents/dw-incidents/ DESCRIPTION: Diagnose data incidents, trace root cause through the lineage graph, and run remediation playbooks — with humans in the loop for anything novel. ============================== The Incidents agent is the on-call responder for your data platform. Feed it anomaly signals — a metric that moved, a pipeline that failed, an SLA that slipped — and it classifies the incident into one of six types (schema change, source delay, resource exhaustion, code regression, infrastructure, quality degradation), assigns a severity, and suggests remediation. It works the way a good SRE does: timeline first, then competing hypotheses, then evidence, then escalation. For root cause, it doesn't guess from the symptom. It traverses the lineage graph upstream, queries execution logs, and cross-references incident history to build a causal chain with confidence scores — distinguishing the symptom (a missing row count) from the cause (a late-arriving upstream batch, a schema change that broke a join). ## Key capabilities - **Incident classification.** `diagnose_incident` takes raw anomaly signals and returns an incident type, severity, and suggested remediation actions. - **Graph-based root cause analysis.** `get_root_cause` traverses lineage up to five or more hops upstream and returns a causal chain with confidence scores, not a single guess. - **Playbook remediation with a confidence gate.** `remediate` runs known playbooks — restart task, scale compute, apply schema migration, switch backup source, backfill data. Automatic execution requires diagnosis confidence above 0.95; anything novel generates a diagnosis report and routes to a human for approval. - **Dry-run mode.** Simulate a remediation before executing it to see exactly what would happen. - **Incident memory.** `get_incident_history` uses vector similarity to find past incidents like the current one, so the third occurrence of a pattern is diagnosed faster than the first. - **Structured incident communication.** Output leads with severity, phase, and impact surface — which assets, which downstream consumers, since when — in scannable form. ## Example prompts > "Row counts on `analytics.daily_orders` dropped 40% overnight — diagnose it." > "What's the root cause of this incident? Walk the lineage upstream and show me the causal chain." > "Have we seen an incident like this before? Show me similar past incidents." > "Dry-run the backfill playbook for this incident before we execute anything." ## Connect it to your stack - **Warehouses and lakehouses** — Snowflake, BigQuery, Databricks for the signals and the affected assets. - **Orchestration** — Airflow and friends, for pipeline run context and restart playbooks. - **Alerting** — PagerDuty, Slack, Microsoft Teams, OpsGenie for paging and resolution. See the [connector catalog](/connectors/) for setup. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data: a realistic set of incidents, metrics, and lineage you can diagnose end to end before any credential exists. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - In 🟡 Evaluation, diagnosis and remediation run against the sample estate — no real system is touched until connections are verified. - Auto-remediation is deliberately narrow: only known patterns above the 0.95 confidence threshold execute without a human. Everything else stops at a diagnosis report and an approval request. That's a floor, not a temporary restriction. - Remediation playbooks are a fixed set; incidents outside them route to a human with the evidence gathered so far rather than an improvised fix. - Root cause quality depends on lineage coverage. If an upstream system isn't connected, the causal chain stops at the boundary and says so. ============================== PAGE: Insights (dw-insights) URL: https://docs.dataworkers.io/products/agents/dw-insights/ DESCRIPTION: Natural-language analytics against a governed warehouse — with pre-registered analysis plans, reproduced numbers, and anomalies root-caused before they become conclusions. ============================== The Insights agent answers business questions against your warehouse in plain language: it translates the question to SQL, executes it read-only, and returns structured results with a narrative — key findings, trends, and recommended next steps. It supports conversational follow-ups within a session, so "now break that down by region" works without restating the question. What separates it from a generic NL-to-SQL tool is discipline. The agent works like a senior data scientist who has been burned by a dashboard that lied: it pins the metric's definition, grain, and time window before querying, resolves metric definitions against the catalog rather than re-deriving them, and pre-registers confirmatory analyses before slicing. A surprising number gets root-caused — denominator change, join fan-out, late-arriving data, or real signal — before it's allowed to become a conclusion. It would rather tell you "we can't tell yet" than ship a false driver. ## Key capabilities - **Natural-language queries.** `query_data_nl` turns a business question into SQL, executes it read-only, and returns structured results — with the generated SQL visible for inspection, and session context for follow-ups. - **Insight narratives.** `generate_insight` analyzes query results into findings, trends, and actionable recommendations — shipped only after the number is reproduced. - **Anomaly explanation.** `explain_anomaly` produces a business-language explanation of a metric anomaly, with possible causes and suggested actions. - **Pre-registered analysis plans.** `register_analysis_plan` freezes the hypothesis, primary metric, and planned cuts before the first confirmatory query; cuts discovered mid-analysis are labeled exploratory until re-tested. - **Alerts and schedules.** `create_alert` and `schedule_insight` keep a verified metric watched, so a regression doesn't go unnoticed; `export_insight` ships the finding. - **Catalog-resolved metrics.** Metric definitions come from the [Catalog & Context agent](/products/agents/dw-context-catalog/), so your number matches the rest of the company's. ## Example prompts > "Why is weekly activation down? Define the metric first, then show me the decomposition." > "Pre-register this analysis: hypothesis, primary metric, and the three cuts we agreed on." > "Revenue per account spiked 18% this week — explain the anomaly before we celebrate." > "Turn yesterday's churn query into a weekly scheduled insight with an alert on regression." ## Connect it to your stack - **Warehouses** — [Snowflake](/connectors/snowflake/), [BigQuery](/connectors/bigquery/), [Databricks](/connectors/databricks/): where the questions get answered. - **Catalog and semantic context** — [dbt](/connectors/dbt/), [DataHub](/connectors/datahub/), and the rest of the [connector catalog](/connectors/) supply the metric definitions and lineage the rigor depends on. ## Works before you connect anything The agent starts in 🟡 Evaluation on a built-in sample estate — you can ask questions, generate insights, and walk the full pre-registration workflow before any credential exists. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - **Read-only, by design.** `query_data_nl` executes read-only queries against governed, modeled tables. This agent never mutates your warehouse. - NL-to-SQL is generation, not magic: the agent inspects the generated SQL for join grain and grouping before trusting the rows, and you can too — the SQL is always shown. - The rigor rules cut both ways: a cut you didn't pre-register comes back labeled exploratory, and a number without a reproduced query and an interval won't be reported as a finding. If you want fast unlabeled slices, this agent will push back. - Answer quality is bounded by catalog coverage: where no canonical metric definition exists, the agent states the definition it chose rather than pretending there was one. ============================== PAGE: Migration (dw-migration) URL: https://docs.dataworkers.io/products/agents/dw-migration/ DESCRIPTION: Oracle, Teradata, and Redshift to Snowflake migration with a self-correcting translate → judge → parity loop and a completion gate that fails closed. ============================== The Migration agent plans, translates, and verifies warehouse migrations. Its core is rule-based SQL dialect translation — Oracle, Teradata, and Redshift to Snowflake, with MySQL and PostgreSQL rules as well — covering type mappings, `DECODE` conversion, and the dialect edge cases that become silent correctness bugs if left as warnings. Around the translator sits the part most migration tools skip: verification. Translations are checked for parity against source data, optionally judged for preserved *intent*, and re-translated when they fail — a self-correcting loop rather than a one-shot converter. The agent thinks like a DBA who has been paged for a botched cutover: dry-run before execute, parity before cutover, rollback plan before forward plan. Irreversible steps are flagged and require human approval before execution — that is the default, not an option. ## Key capabilities - **Migration assessment.** `assess_migration` inventories objects in the source system and classifies migration complexity, including PII columns and an effort estimate. - **SQL dialect translation.** `translate_sql` and `batch_translate_sql` apply rule-based mappings for known patterns; complex SQL can fall back to your model. - **The self-correcting loop.** `migrate_with_validation` runs translate → (judge) → parity → re-translate until parity passes or the round limit is hit. On non-convergence it returns the best attempt with review annotations for a human — never a silent pass. - **Intent judging.** `judge_translation` checks whether a translation preserves the *meaning* of the source, not just matching rows on a sample — catching dropped `DECODE` defaults, changed NULL handling, and silent join-type changes. - **Parity validation.** `validate_migration` and `run_parallel_comparison` compare row counts, column stats, and sample hashes between source and target, and report divergences and match rates. - **Dependency and blast-radius mapping.** `map_migration_dependencies` builds the downstream-consumer graph and returns a wave-ordered migration plan (topological sort with cycle detection), flagging high-blast-radius objects. - **A completion gate that fails closed.** `gate_migration_complete` refuses to certify a migration done until validated and parity-checked coverage clears the threshold. An empty or under-covered scope is "not done." - **Fleet status.** `get_migration_status` gives per-run progress with per-object drill-down: pending, translating, judging, validating, needs review, converged, failed. ## Example prompts > "Assess our Teradata warehouse for migration to Snowflake — inventory, complexity, effort." > "Translate this Oracle procedure to Snowflake and judge whether the translation preserves intent." > "Run the self-correcting migration on these 40 views and show me anything that didn't converge." > "Map the dependency graph for `oracle-prod-1` and give me a wave-ordered plan with blast radius." > "Is this migration actually done? Gate it at 95% parity coverage." ## Connect it to your stack - **[Snowflake](/connectors/snowflake/)** — the migration target for parity checks against real data. - **Catalog connectors** — [DataHub](/connectors/datahub/), [dbt](/connectors/dbt/), and the rest of the [connector catalog](/connectors/) enrich dependency mapping with lineage, so blast radius includes consumers outside the source inventory. - The translator itself needs no connection — it works on SQL you paste in. ## Works before you connect anything The agent starts in 🟡 Evaluation, and the translation rules are real logic even there: `translate_sql` works on *your* SQL from the first minute, no credential required. Parity checks and inventory run against built-in sample data until systems are connected and verified. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - The judge **fails open**: if the judge model is unavailable or returns unparseable output, the verdict is `pass`. Treat the judge as an extra tripwire, not the last line of defense — parity validation is the harder check. The judge requires your own model key (bring-your-own-model, like everything in the swarm). - The completion gate **fails closed** — deliberately. Expect it to say "not done" until you have real coverage, including on an empty scope. - The self-correcting loop is bounded by `maxRounds`. When it doesn't converge, you get the best attempt annotated for human review, not a certified translation. - Lineage enrichment for dependency mapping degrades gracefully: if no catalog is connected, the plan is built from the source inventory only, and the result says so. - Dialect edge cases flagged during translation are treated as blockers, not warnings — the agent will stop and tell you rather than guess. ============================== PAGE: MLOps & Models (dw-ml) URL: https://docs.dataworkers.io/products/agents/dw-ml/ DESCRIPTION: Experiment tracking, a staged model registry, feature pipelines, drift detection, explainability, and A/B testing — the ML lifecycle as an agent. ============================== The MLOps & Models agent manages the machine-learning lifecycle end to end: experiment tracking, a versioned model registry with stage management, feature pipelines with drift statistics, model explainability, and guarded rollout with A/B testing. It treats a model as one service in a graph — data, features, training, registry, serving, monitoring — not a notebook artifact, and it treats a feature as a versioned product with an entity key, null semantics, and a freshness expectation, not a column name in a file. Its engineering instincts are the ones that survive production: assume train–serve skew until proven otherwise, audit for leakage before believing any offline gain, require a reproducible run before promotion, and prove the rollback path before any traffic shifts. A "better model" with no kill switch is, in this agent's view, not shippable. ## Key capabilities - **Experiment tracking.** `create_experiment`, `log_metrics`, and `compare_experiments` give you MLflow-compatible run tracking with step-based training curves and cross-run comparison including significance tests. - **Model registry with stages.** `register_model` and `get_model_versions` version models through development, staging, and production, with stage history and lineage. - **Budget-aware AutoML training.** `ml_automl_train` is the real-compute path: you set a time budget, not an algorithm, and a FLAML-based search returns the best estimator, tuned hyperparameters, and real cross-validation scores, persisting a versioned model artifact. - **Feature pipelines and stats.** `create_feature_pipeline` generates scheduled feature jobs (sinks include Snowflake, BigQuery, Redis); `get_feature_stats` reports distributions, null rates, and drift scores against training baselines. - **Explainability.** `explain_model` produces SHAP-style feature-importance reports (permutation and gain methods also available), down to individual-prediction explanations. Today these are deterministic simulations for evaluating the workflow — not values computed against your live model. - **Drift detection.** `detect_model_drift` separates data drift from concept drift using KS, PSI, and Chi-squared tests with configurable thresholds. - **Guarded rollout and A/B testing.** `deploy_model` supports canary, shadow, and blue-green strategies; `ab_test_models` configures traffic splits and reports metric comparison with statistical significance. - **Feature suggestions and model selection.** `suggest_features` ranks feature-engineering candidates from a dataset profile; `select_model` recommends architectures for the problem type and constraints. ## Example prompts > "Create an experiment for the churn model and run an AutoML search with a 20-minute budget." > "Compare the last three experiment runs — which wins on calibration and the worst slice, not just AUC?" > "Set up a feature pipeline for these five features with a daily refresh into Snowflake." > "Has the fraud model drifted in the last 7 days? Separate data drift from concept drift." > "A/B test v12 against the champion at a 10% split, primary metric precision@k." ## Connect it to your stack - **Warehouses and feature sinks** — [Snowflake](/connectors/snowflake/), [BigQuery](/connectors/bigquery/), [Databricks](/connectors/databricks/) for training data and feature storage. - **ML tooling** — MLflow and Weights & Biases are in the ML category of the [connector catalog](/connectors/). - **Catalog** — [DataHub](/connectors/datahub/) and [dbt](/connectors/dbt/) supply the feature and metric definitions the agent resolves instead of re-deriving. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data — you can walk the full lifecycle, from experiment to registry to a simulated rollout, before any credential exists. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - `train_model` is a deterministic simulation kept for compatibility — for a real fit, use the AutoML path (`ml_automl_train`). That path requires a Python runtime with FLAML available; without one it fails with a clear error rather than faking metrics, and a timed-out search is reported as a failed run, not a shippable best-so-far. - Training budgets are time and trials — never dollars. - In 🟡 Evaluation, registry, deployment, and monitoring act on the sample environment; nothing reaches real serving infrastructure until connections are verified. - The agent's own discipline is a limit you'll feel: no promotion without a reproducible run, and no rollout without a tested rollback path. It will refuse shortcuts rather than register an unreproducible champion. - Data-quality breaks it finds are handed to the [Quality agent](/products/agents/dw-quality/) rather than silently worked around. ============================== PAGE: Pipelines & Ingestion (dw-pipelines) URL: https://docs.dataworkers.io/products/agents/dw-pipelines/ DESCRIPTION: Turn a natural-language description into a validated, deployable data pipeline — with templates, sandbox validation, and Airflow deployment. ============================== The Pipelines & Ingestion agent turns a plain-English description of what you want moved and transformed into a working pipeline. It decomposes your description into extraction, transformation, loading, testing, and deployment tasks, generates the code (SQL, Python, or dbt), and targets the orchestrator you already run — Airflow, Dagster, or Prefect. It also covers the ingestion side: EL, CDC, and replication patterns, including Iceberg `MERGE INTO`. It doesn't work alone. When it generates a pipeline, it registers the new asset in the catalog, asks the Quality agent to create quality tests for it, and checks schema compatibility — so a pipeline born here arrives with context, tests, and lineage instead of as an orphan. ## Key capabilities - **Natural-language pipeline generation.** `generate_pipeline` decomposes a description into extract/transform/load/test/deploy tasks and generates the code, with template fallback when generation isn't confident. - **Real validation before anything ships.** `validate_pipeline` runs sandbox execution: Python AST parsing, SQL syntax checking, and YAML schema validation, plus semantic-layer validation when the catalog agent is reachable. - **Airflow deployment with verification.** `deploy_pipeline` writes DAG files via filesystem, S3, or git-sync and verifies the deployment through the Airflow REST API — it doesn't just drop a file and hope. - **Versioned specs in Git.** Deployment can commit the pipeline specification as YAML to your repo, so every deployed pipeline has a reviewable history. - **A template library for common patterns.** `list_pipeline_templates` covers ETL, ELT, CDC, streaming, reverse-ETL, and data-quality patterns, filterable by orchestrator — use one as a starting point instead of generating from scratch. - **Cross-agent registration.** Generated pipelines are registered in the catalog, get quality tests created, and are checked for schema compatibility automatically. ## Example prompts > "Build a daily pipeline that loads new orders from Postgres into Snowflake, deduplicates > on order_id, and merges into the `analytics.orders` Iceberg table." > "Validate this pipeline spec before I deploy it — check the SQL and the Python." > "What CDC templates do you have for Airflow?" > "Deploy the validated orders pipeline to staging and commit the spec to the main branch." ## Connect it to your stack - **Orchestration** — Airflow (deployment and verification), plus Dagster and Prefect as generation targets. - **Warehouses and lakehouses** — Snowflake, BigQuery, Databricks as pipeline sources and targets. - **dbt** — as a code language for generated transformations. See the [connector catalog](/connectors/) for setup. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data — the generation, templating, and validation logic is the real thing, run against a realistic sample estate. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - In 🟡 Evaluation, `deploy_pipeline` records the deployment locally — nothing reaches a real orchestrator until Airflow is configured and verified. - The verified deployment path is Airflow today. Dagster and Prefect are supported as generation targets; deployment to them is not yet implemented. - Semantic-layer validation only runs when the Catalog & Context agent is reachable — the syntax and sandbox checks still run without it. - Deploying is a governed write: it follows the propose–approve–receipt path like every other write in the swarm. ============================== PAGE: Quality (dw-quality) URL: https://docs.dataworkers.io/products/agents/dw-quality/ DESCRIPTION: Five-dimension quality scoring, statistical anomaly detection against 14-day baselines, and SLAs that alert with evidence instead of noise. ============================== The Quality agent monitors your data across five dimensions — completeness, accuracy, consistency, freshness, and uniqueness — and rolls them into a weighted 0–100 score per dataset, with the breakdown and trend visible so you know *why* a score moved. Anomaly detection is statistical (z-score against 14-day baselines), and detected anomalies are deduplicated hard: 50–100 raw signals typically collapse to 5–10 actionable ones, because an alert feed nobody reads is worse than no alert feed. Its operating discipline matters as much as its math. It re-verifies a failure before alerting on it, distinguishes "the data is wrong" from "the check is wrong", and never silently corrects, filters, or imputes data to make a check pass — any correction goes through the governed write path with human approval. Genuine breaks are handed to the Incidents agent rather than absorbed into monitoring. ## Key capabilities - **Full quality profiling.** `run_quality_check` profiles null rates, uniqueness, distributions, referential integrity, freshness, and volume for a dataset — all columns or a chosen subset. - **A score you can interrogate.** `get_quality_score` returns the 0–100 score with its per-dimension breakdown and trend, not just a number. - **Deduplicated anomaly detection.** `get_anomalies` lists detected anomalies classified by severity (critical / warning / info), deduplicated to the actionable set by default. - **Quality SLAs.** `set_sla` defines metric thresholds with severity levels per dataset; violations are designed to trigger alerts within a minute. - **Tests born with the pipeline.** `create_quality_tests_for_pipeline` generates quality test specs for a pipeline — this is what the Pipelines agent calls so new pipelines arrive with tests. - **Estate-level summary.** `get_quality_summary` aggregates quality across datasets so you can see where the estate stands, not just one table. ## Example prompts > "Run a quality check on `analytics.orders` — all columns." > "Why did the quality score on the customers table drop this week?" > "Show me critical anomalies from the last 48 hours — deduplicated, not the raw feed." > "Set an SLA on `finance.revenue_daily`: null rate under 1%, freshness under 6 hours, > critical severity." ## Connect it to your stack - **Warehouses and lakehouses** — Snowflake, BigQuery, Databricks for profiling real tables. - **Quality suites** — Great Expectations, Soda, Monte Carlo results flow in through the connector gateway. - **dbt** — test results as a quality signal. See the [connector catalog](/connectors/) for setup. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data — the profiling, scoring, and z-score anomaly algorithms are the real thing, run on a realistic sample estate. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - In 🟡 Evaluation, scores and anomalies describe the sample estate, not your warehouse — useful for judging the workflow, meaningless as a statement about your data. - Anomaly detection needs a baseline: after connecting a new system, expect the 14-day baseline to build before statistical detection is at full strength. - The agent will not fix data for you silently. Corrections are explicit, traced, and human-approved — if you want a number quietly adjusted, this is the wrong tool. - A failed check is a signal, not an alert. The agent re-verifies before alerting, which trades a little latency for a lot less noise. ============================== PAGE: Schema (dw-schema) URL: https://docs.dataworkers.io/products/agents/dw-schema/ DESCRIPTION: Detect schema changes, classify them as breaking or not, assess blast radius through lineage, and generate migrations with rollback built in. ============================== The Schema agent watches the shape of your data. It detects schema changes by monitoring `INFORMATION_SCHEMA`, schema registries, and Git webhooks, and classifies each change as breaking or non-breaking — including rename detection, so a renamed column isn't misdiagnosed as a drop plus an add. Snapshots let it track how a schema evolved over time, not just what it looks like now. When a change lands (or before you make one), it answers the question that actually matters: what breaks? It traverses the lineage graph to identify every affected pipeline, view, dashboard, ML model, and API, then generates a migration to move forward safely — with the rollback script written at the same time as the forward one. ## Key capabilities - **Change detection with classification.** `detect_schema_change` monitors `INFORMATION_SCHEMA`, schema registries, and Git webhooks, and labels each change breaking or non-breaking. Scan a single table or a whole source. - **Impact assessment through lineage.** `assess_impact` walks the lineage graph and lists affected pipelines, views, dashboards, models, and APIs before you commit to a change. - **Migration generation with rollback.** `generate_migration` produces forward SQL, rollback SQL, and updates for affected systems (SQL, dbt, API), validated with sqlglot before you see it. - **Safe application.** `apply_migration` supports blue/green and rolling strategies, a `dryRun` mode that validates without executing, and automatic rollback capability — downstream agents are notified when a migration lands. - **Compatibility checks.** Real Avro/JSON schema compatibility rules catch a change that would break consumers at the contract level, not just the table level. - **Snapshot-based evolution.** Schema snapshots and a change log let you ask what a table looked like before the incident, and what changed since. ## Example prompts > "Did anything change in `prod.core` this week? Flag anything breaking." > "If I drop `orders.legacy_status`, what breaks downstream?" > "Generate a migration for this column type change — and the rollback." > "Dry-run this migration against staging before we apply it." > "Is this new event schema backward-compatible with what consumers expect?" ## Connect it to your stack - **Warehouses and lakehouses** — Snowflake, BigQuery, Databricks, and other `INFORMATION_SCHEMA` sources; Iceberg snapshot-based evolution. - **Streaming** — Kafka Schema Registry for contract-level compatibility. - **dbt** — generated migrations can target dbt models alongside raw SQL. See the [connector catalog](/connectors/) for setup. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample schemas — the diff, impact-traversal, DDL-generation, and compatibility logic is the real thing, run on a sample estate. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - In 🟡 Evaluation, `apply_migration` records the migration locally — no real database is altered until a connection is verified. - Applying a migration is a governed write: proposed, approved, executed with a receipt. Irreversible changes always ask first. - Impact assessment is only as complete as your lineage. Systems that aren't connected are invisible to the blast-radius calculation — the result tells you what it could see, not what it couldn't. ============================== PAGE: Data & Cloud Security (dw-security) URL: https://docs.dataworkers.io/products/agents/dw-security/ DESCRIPTION: DSPM and CSPM in one agent — discover and classify sensitive data, scan cloud posture, and route ranked findings to the agent that can fix them. ============================== The Data & Cloud Security agent unifies two jobs that usually live in two consoles: data security posture (DSPM) — discovering and classifying sensitive data such as PII, PHI, PCI, and secrets, and scoring the risk it carries — and cloud security posture (CSPM) — scanning for misconfiguration, public exposure, weak encryption, exposed secrets, and over-privileged identities. What makes it an agent rather than another scanner is what happens after a finding. It ranks findings and routes each one to the agent that can act: remediation such as an access revocation goes to the [Data Access & Governance agent](/products/agents/dw-governance/), and active threats go to the [Incidents agent](/products/agents/dw-incidents/). You ask a question in plain language; you get a ranked answer and a path to the fix. ## Key capabilities - **Sensitive-data discovery and classification.** Finds where PII, PHI, PCI data, and secrets live in your estate and classifies what it found. - **Data risk scoring.** Scores the risk each finding carries, so the queue starts with what matters rather than a flat list. - **Cloud misconfiguration and public-exposure scanning.** Checks posture for misconfiguration, public exposure, and weak encryption. - **Exposed-secrets and over-privilege detection.** Flags exposed secrets and over-privileged identities as posture findings. - **Ranked findings, not raw dumps.** Findings come back ordered by risk, in a form you can act on from your terminal. - **Routing to the fix.** Remediation routes to Data Access & Governance (for example, a revoke); active threats route to Incidents. The security agent hands off; the receiving agent's governed write path applies. - **AWS Security Hub backend.** Cloud posture findings can be pulled from a connected AWS Security Hub — the one live posture backend wired today. ## Example prompts > "Where does PII live in my estate, and which of it is highest risk?" > "Scan my cloud posture for public exposure and weak encryption, ranked by severity." > "Are any identities over-privileged for the data they touch?" > "Take the top finding and route it — who fixes it, and what's the proposed action?" ## Connect it to your stack - **AWS Security Hub** — the live backend for cloud posture findings. - **Warehouses and catalogs** — [Snowflake](/connectors/snowflake/), [BigQuery](/connectors/bigquery/), [Databricks](/connectors/databricks/) and the other catalog connectors give the swarm the estate context that findings are ranked and routed against. - **Identity and alerting systems** are covered in the [connector catalog](/connectors/). ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample data — a realistic estate with plausible findings you can discover, rank, and route end to end before any credential exists. It earns 🟢 Connected per system through a passing live test. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - **What this agent is:** agentic unification and routing of DSPM + CSPM findings. **What it is not:** a replacement for a dedicated scanner — we do not claim scan parity with tools like Wiz or BigID. If you run one of those, this agent complements it by putting findings where your agents already work. - **It surfaces findings; it does not enforce in production.** Fixes go through the receiving agent's governed, approval-gated write path — the security agent itself never mutates your systems. - One live posture backend is wired today (AWS Security Hub). The other capabilities run on sample data until their systems are connected and verified. - Routing quality depends on which agents and systems are connected: a revoke can only be proposed where Data Access & Governance has a connected target. ============================== PAGE: Streaming (dw-streaming) URL: https://docs.dataworkers.io/products/agents/dw-streaming/ DESCRIPTION: Kafka Connect configuration generation, consumer-lag monitoring, stream health, and tuning recommendations that are always human-applied. ============================== The Streaming agent handles the operational core of Kafka-based pipelines: getting a stream topology configured correctly, and knowing when it's falling behind. Describe a source and a sink — say, a Debezium Postgres source into a Snowflake sink — and it validates the topology and generates the Kafka Connect JSON configuration, converter classes included, with your partition count, replication factor, and SLA targets baked in. Once a stream runs, the agent watches it: consumer lag per partition, connector status, latency, and error rates, aggregated into per-connector and overall stream health. When something drifts — lag climbing, throughput sagging against an SLA — it generates tuning recommendations. It recommends; you apply. Nothing is auto-tuned behind your back. ## Key capabilities - **Kafka Connect config generation.** `configure_stream` validates a stream topology and produces working Kafka Connect JSON for source and sink connectors — the real config generation logic, including converter classes. - **SLA-aware setup.** Pass SLA targets (max latency, max lag in records) at configuration time and they become the baseline the tuning engine judges against. - **Consumer-lag monitoring.** `monitor_lag` returns lag per partition for a topic, plus total lag across partitions — the first number you want when a downstream table goes stale. - **Stream health aggregation.** `get_stream_health` rolls connector status, latency, and error rates into per-connector and overall health for a topology. - **Tuning recommendations.** `get_recommendations` analyzes lag, throughput, and health against your SLAs and proposes concrete tuning changes — human-applied, never auto-applied. ## Example prompts > "Generate a Kafka Connect config for a Debezium Postgres source into a Snowflake sink on topic `orders-cdc`, 12 partitions." > "What's the consumer lag on `orders-cdc`, per partition?" > "Give me the health of the `orders-cdc` topology — connectors, latency, error rates." > "Lag has been climbing all afternoon. What tuning changes do you recommend?" ## Connect it to your stack - **Kafka Schema Registry** — the streaming entry in the [connector catalog](/connectors/); schema registration and compatibility checks flow through it. - **Warehouse sinks** — configs commonly target sinks like [Snowflake](/connectors/snowflake/); connect the warehouse to verify the receiving end. ## Works before you connect anything The agent starts in 🟡 Evaluation, and config generation is real logic even there: `configure_stream` produces valid Kafka Connect JSON for your topology before any credential exists. Lag and health monitoring run against built-in sample data until your systems are connected and verified. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - This is a deliberately small surface — configuration, lag, health, tuning. It manages Kafka Connect topologies; it is not a stream processor and doesn't run your Flink or Spark jobs. - Recommendations are never auto-applied. If you want a hands-off tuner, this agent will disappoint you on purpose: it proposes, you decide. - In 🟡 Evaluation, lag and health numbers describe the sample topology, not your cluster. Generated configs are real, but validate them against your environment before deploying. - Monitoring depth depends on what's connected; where a metric isn't available from your setup, the agent says so rather than inventing a number. ============================== PAGE: Usage Intelligence (dw-usage-intelligence) URL: https://docs.dataworkers.io/products/agents/dw-usage-intelligence/ DESCRIPTION: Zero-LLM analytics on how practitioners actually use the platform — adoption, sessions, workflow patterns, heatmaps, and a hash-chained activity log. ============================== The Usage Intelligence agent answers the question every platform owner eventually asks: is anyone actually using this, and how? It measures practitioner interactions with the Data Workers platform itself — which tools and agents are called, by whom, when, in what sequences — and turns that into adoption dashboards, session analytics, usage heatmaps, and anomaly detection. It is deliberately **zero-LLM**: every number comes from deterministic computation over recorded activity, not from a model's summary. That makes it cheap to run continuously and makes its outputs reproducible — the same activity log always yields the same dashboard. It's how you find out which agents earned adoption, which are shelfware, and where a sudden usage drop is telling you something broke. ## Key capabilities - **Tool and agent usage metrics.** `get_tool_usage_metrics` reports usage volume, unique users, trend direction, and response times, grouped by tool, agent, or user. - **Adoption dashboards.** `get_adoption_dashboard` classifies agents and tools as adopted, growing, underused, or shelfware, against a configurable adoption threshold. - **Session analytics.** `get_session_analytics` measures session duration, depth (tools per session), agents per session, and classifies users as power, regular, or occasional. - **Workflow patterns.** `get_workflow_patterns` finds the multi-tool and multi-agent sequences practitioners actually chain together, and how much usage is standalone versus part of a workflow. - **Usage heatmaps.** `get_usage_heatmap` shows when and where people interact — hourly, daily, or agent-by-user. - **Usage anomaly detection.** `detect_usage_anomalies` flags sudden drops (friction), unusual spikes (automation loops or incidents), and behavior shifts, at configurable sensitivity. - **Hash-chained activity log.** `get_usage_activity_log` retrieves the SHA-256 hash-chained record of who called which tool, when, with what outcome — and verifies chain integrity. - **Agent health.** `list_active_agents` and `check_agent_health` report which agents are running and how they're doing. ## Example prompts > "Which agents got adopted last month and which are shelfware?" > "Show me the usage heatmap for the last 30 days — when do people actually use this?" > "What tool sequences do our power users chain together?" > "Usage of the quality agent dropped 60% this week — is that friction or an outage?" > "Pull the activity log for last week and verify the hash chain." ## Connect it to your stack Usage Intelligence measures the swarm itself, so it lights up as a side effect of using the platform rather than through its own warehouse credentials. The more agents and [connectors](/connectors/) your team runs, the more its dashboards have to say. ## Works before you connect anything The agent starts in 🟡 Evaluation on built-in sample activity — a simulated 30 days of usage — so you can explore every dashboard, pattern, and heatmap before your own history exists. See [Verify your setup](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/). ## Limits, honestly - Scope is the Data Workers platform: this is analytics on practitioner interaction with the swarm, not a general product-analytics tool for your applications. - In 🟡 Evaluation the numbers describe the sample activity, not your team. Real dashboards need real usage history, which accumulates only after your team starts working through the platform. - Zero-LLM means deterministic, but also literal: it reports what happened, not why. Pair a surprising pattern with the relevant specialist agent for interpretation. - Anomaly detection is statistical; at high sensitivity expect some false positives, and tune the sensitivity down if the noise outweighs the catches. ============================== PAGE: Conductor URL: https://docs.dataworkers.io/products/conductor/ DESCRIPTION: The always-on operations loop for your data estate — propose-first, approval-gated, in preview. ============================== :::caution[Preview] The Conductor is in preview. Its guardrails are the load-bearing feature and they ship first: it proposes rather than silently acts, and irreversible actions always require your approval — enforced in code, not judgment. ::: The Conductor is the always-on layer of the swarm: instead of waiting for a prompt, it watches your connected estate — documentation drift, reversible hygiene, quality baselines, cost anomalies — and compiles standing goals ("tables in `analytics` stay documented") into routine work for the fleet. ## The two modes You choose how much initiative it gets, and you can switch anytime: - **Approve-edits** — the Conductor proposes every change; you click approve. Start here, watch it for a day, then flip to Auto in one click. No penalty, no re-onboarding. - **Auto** — it performs *reversible* upkeep itself and tells you after; stop or undo anytime. ## The floor that never moves Regardless of mode: **reversible** means you can undo it — a description, a tag, a re-run. **Irreversible** means you can't — a delete, an access grant, anything that spends money. Irreversible actions always ask first, in every mode, with the exact change and blast radius shown. Even with no zones configured, the Conductor can never spend, change access, or delete without asking. Everything it does lands in the same receipt-and-audit trail as any governed write, and its track record — changes made, kept, undone — is visible, so trust is earned from your own history rather than asserted. ## Where it runs Hosted by us as part of your workspace on every plan — inside your own deployment boundary on Enterprise dedicated/on-prem. Enabling it and choosing its mode is part of [pilot onboarding](https://docs.dataworkers.io/onboarding/getting-started/your-pilot/). Self-hosting the Conductor on your own infrastructure is also supported for teams that require it — raise it with your onboarding engineer. ============================== PAGE: Data Workers Data Context Wizard URL: https://docs.dataworkers.io/products/data-context-wizard/ DESCRIPTION: The governed, federated knowledge graph your agents read, reason over, and write back to — context your agents can trust. ============================== The Data Workers Data Context Wizard is the semantic and knowledge graph for your data platforms — the governed, federated context graph an enterprise's data agents read, reason over, and write back to, so they stop answering confidently and wrong. **Tagline, and the actual design goal: *context your agents can trust.*** ## Why a context graph, not another catalog A catalog is where you *put* context — and it goes stale between crawls. The Context Wizard is the layer a governed fleet keeps alive: agents harvest context from the systems you connect, write back what they learn under governance, and every fact carries provenance and confidence. One agent can edit a graph; a governed fleet can *compound* one — that's the difference. Three things make the write-back safe enough to trust: - **Provenance and confidence on every fact.** Nothing enters the graph anonymous. - **Governed promotion.** A fact only becomes *authoritative* when a named human approves it — promotion is a gate, not a default. - **Isolation and audit.** Tenant-scoped, with hash-chained audit on governed writes. ## The graph is federated The Context Wizard reads across the platforms you already run — warehouses, catalogs, dbt, lineage systems — through the same [connector set](/connectors/) as the rest of the platform, rather than requiring your estate to move into one vendor's catalog first. ## What's included - **Every plan (pilot, Scale, Enterprise):** the context graph the agents read and write back to under governance, and the **authoring surface** — your team curates definitions, business rules, and ownership in the browser, runs the promotion workflow, and manages the graph's warmth. - **Enterprise adds:** compliance export of the graph and audit chain, inside your deployment boundary. ## Getting started The graph starts cold and warms in phases — harvest, reconcile, curate, promote. The practical guide, including what's automated versus what only you can provide, is [Warming up your knowledge graph](/products/warming-your-graph/). Two honest notes before you start: :::note[Scope your first harvest] Start with one schema that matters, not the whole estate. Bulk harvest of very large estates is still being hardened — a scoped start gives you trustworthy results and a clean expansion path. ::: :::note[Self-maintaining is enabled separately] The always-fresh, self-maintaining graph — agents re-harvesting on a standing schedule — is enabled per deployment rather than on by default. Until it's enabled for your workspace, treat the graph as current-as-of-last-harvest. ::: ============================== PAGE: Deployment options URL: https://docs.dataworkers.io/products/deployment/ DESCRIPTION: Self-hosting the Data Workers agent swarm — npx, source, Docker Compose, Kubernetes, and air-gapped installs. ============================== Every plan level can self-host the agent core. This page goes from lightest to heaviest. ## Zero-install (recommended start) No infrastructure at all — agents run in-process with in-memory adapters, ideal for evaluation and for individual engineers: ```bash claude mcp add data-workers -- npx -y dw-claw ``` In-memory adapters are suitable for development and evaluation; production deployments should use real infrastructure adapters (below). ## From source ```bash git clone https://github.com/DhanushAShetty/data-workers.git cd data-workers npm install ./start-agent.sh dw-connectors # any agent by name ``` Agents run from source via `tsx` — no build step. Cross-platform (Windows-safe) equivalents: ```bash node plugin/scripts/install.mjs # prereq check + install node plugin/scripts/start-agent.mjs ``` ## Docker Single image: ```bash docker pull ghcr.io/dhanushashetty/data-workers:latest docker run -it ghcr.io/dhanushashetty/data-workers:latest ``` Full stack with real infrastructure via Compose: ```bash docker compose -f docker-compose.yml -f docker/docker-compose.agents.yml --env-file .env up -d ``` The production profile expects PostgreSQL 16 (with pgvector) and Redis 7, plus one LLM provider key; Neo4j 5, Kafka, and Airflow are optional, enabled when the matching agents need them. **Requirements:** Node.js ≥ 20, npm ≥ 10, Docker ≥ 24 (Kubernetes ≥ 1.28 for the manifests below). ## Kubernetes Example manifests — per-agent Deployments plus the shared services — are in the repo's deployment guide (`DEPLOYMENT.md`). Sizing guidance: agents are I/O-bound; start small and scale replicas on the agents your team actually uses (usage intelligence will tell you which those are). ## Air-gapped (Enterprise) Fully offline installs are supported and documented in the repo's air-gap deployment guide (`docs/DEPLOYMENT-AIRGAP.md`): mirrored registries for npm and OCI images, offline licensing, and local model inference (Ollama/vLLM) so **zero external network calls** leave the boundary — including model calls. Air-gap installs are set up with our team as part of [Enterprise onboarding](https://docs.dataworkers.io/onboarding/getting-started/enterprise/). ## Which one am I supposed to pick? | Situation | Path | |---|---| | Evaluating, or an individual engineer | Zero-install `npx` | | A team rollout | Zero-install per engineer + shared service credentials | | You want to modify agents | From source | | Production self-hosting with persistence | Docker Compose or Kubernetes | | Regulated / offline environment | Air-gapped, with us | ============================== PAGE: Integrations & SDKs URL: https://docs.dataworkers.io/products/integrations/ DESCRIPTION: ChatGPT apps, agent-to-agent protocol, framework adapters, the Python SDK, and the catalog provider SDK. ============================== The swarm speaks MCP natively, and a set of adapters carries it everywhere else. ## ChatGPT Data Workers ships as ChatGPT apps — an All-in-One app plus focused Catalog, Quality, and Incidents apps. Self-hosters can expose an agent to ChatGPT as a custom connector: ```bash docker compose -f docker/docker-compose.chatgpt.yml up # or expose one agent over HTTP: DW_TRANSPORT=http DW_REMOTE_PORT=3001 npx tsx agents/dw-context-catalog/src/index.ts ``` ## Agent-to-agent (A2A) The swarm implements the A2A protocol (spec 0.2.x) so external agent systems can discover and call it: Agent Card discovery, inbound `message/send`, and task status/cancel are shipped. A2A is **opt-in and off by default** — it only listens when you set `DW_A2A_PORT`. Streaming and outbound A2A are later, gated phases. ## Framework adapters For teams orchestrating their own agents: - `@data-workers/a2a-adapter` — A2A server wrapper - `@data-workers/adk-adapter` — Google ADK - `@data-workers/langchain-adapter` — LangChain tools Each wraps the same underlying agents, so capability and governance behavior is identical to the MCP surface. ## Python SDK A typed Python SDK (sync and async clients over the agents) is on the roadmap, not shipped. Until it lands, programmatic access is the MCP protocol itself — any MCP client library can call the agents — or the A2A surface above. [Ask for it](/help/feedback/) if it would unblock you; roadmap priority is demand-driven. ## Build your own connector The `@data-workers/catalog-provider-sdk` is the open interface our own catalog connectors implement — zero runtime dependencies, with a conformance kit to validate your implementation locally. It's the supported path when [your system doesn't have a native connector](/connectors/atlan-alation-collibra/). ## VS Code A VS Code extension ships in the repo (`packages/vscode-extension`) for teams whose engineers live outside the four first-class coding agents. ============================== PAGE: Spellbook Data Catalog URL: https://docs.dataworkers.io/products/spellbook-data-catalog/ DESCRIPTION: The browser control plane over the agent fleet — currently in research preview. ============================== :::caution[Research preview] The Spellbook Data Catalog is in **research preview**. The hosted console is currently offline while we harden it; today it runs as a local preview on your machine, available on the Scale plan. Expect rough edges, and expect this page to change. What's written below describes what works today. ::: The Data Workers Spellbook Data Catalog is the human control plane over the swarm: the browser surface where your team sees what the agents know, watches what they're doing, approves what they propose, and browses the estate the fleet maintains. It is deliberately *not* another system-of-record catalog to maintain by hand. The agents do the cataloging; the Spellbook is where humans stay in command — the receipt surface for the fleet's actions. ## What it does today - **Estate browsing** — tables, schemas, lineage, and quality signals from your connected systems, searchable. - **Connector management** — register a system through a credential form, test it live, see its true 🟡/🟢/🔴 state. A "Paste schema / CSV" path covers systems without a connector. - **Agent activity & approvals** — watch runs, review proposed writes, approve or reject with the change and blast radius shown. - **Onboarding wizard** — a guided first-run flow that gets a first system connected and verified from the browser. ## How it's built (30 seconds, but it matters for security) Split-brain: the UI is a static page; the **brain** — the agents, all credentials — runs on **your machine** via the local launcher on `localhost:4747`. The browser talks to your own machine. We are not a SaaS that ingests your data — no credentials, warehouse data, query results, or PII touch our servers. Tell your security team this first; details in [How your data is handled](https://docs.dataworkers.io/onboarding/security/data-handling/). ## Running it Scale plan, with your customer npm token: ```bash npm config set //registry.npmjs.org/:_authToken= npx @dataworkersproj/agent-console@latest serve --local-ui ``` The full walkthrough is in the [pilot onboarding track](https://docs.dataworkers.io/onboarding/getting-started/your-pilot/). ## What research preview means here - The hosted console site is offline; local preview is the supported path. - Some views run on sample data until their backing system is connected and verified — the console never blends the two silently: sample data is labeled. - Console SSO/SAML/SCIM are not self-service; identity for the browser surface is part of the [Enterprise](https://docs.dataworkers.io/onboarding/getting-started/enterprise/) "set up with us" conversation. - The local launcher needs Node and your npm token, so today a data steward without a dev setup reviews approvals in a shared session with an engineer — a hosted console for non-engineers is exactly what research preview graduates into. - The console's connector card list is a subset of the [full connector set](/connectors/) — anything missing still connects via environment variables in the terminal path. Found something broken or missing? [That feedback](/help/feedback/) directly shapes what graduates out of preview. ============================== PAGE: Warming up your knowledge graph URL: https://docs.dataworkers.io/products/warming-your-graph/ DESCRIPTION: How the Data Context Wizard goes from a cold start to trusted, authoritative context — what's automated and what only you can provide. ============================== Your knowledge graph starts cold. That's not a flaw — it's the honest starting point of any context system that refuses to guess. Warming it up is a sequence of phases, most automated, a few deliberately human. ## The phases | Phase | What happens | Automated, or you? | |---|---|---| | **1 · Harvest** | Agents pull schemas, lineage, query history, and dbt metadata from your connected systems into the graph, every fact stamped with its source. | Automated, once systems are [connected and verified](https://docs.dataworkers.io/onboarding/getting-started/verify-your-setup/) | | **2 · Reconcile** | The same table seen through Snowflake, dbt, and your catalog is merged into one entity; conflicts are surfaced, not silently resolved. | Automated, with conflicts queued for you | | **3 · Enrich** | Agents propose descriptions, ownership guesses, and quality annotations — as *proposals*, with confidence attached. | Automated proposals | | **4 · Curate** | Your team reviews proposals and adds what no system contains: business definitions, the "revenue means ARR for finance" rules, the no-touch lists. | **You** — this is the part only your team knows | | **5 · Promote** | Facts your team stands behind are promoted to *authoritative* — each promotion approved by a named human. | **You**, with the gate enforced by the platform | After warm-up, governed agent write-back keeps the graph improving with use — and the self-maintaining refresh loop, once enabled for your deployment, re-harvests on a standing schedule (see the note on the [Context Wizard page](/products/data-context-wizard/)). ## A practical warm-up plan 1. **Pick one schema that matters.** `analytics` beats `everything` — a scoped harvest completes fast and is checkable by a human who knows the data. 2. **Run the harvest** (Scale: from the Context Wizard surface; Start: agents harvest as they work). Spot-check: do the entity counts look like your estate? If something looks missing, say so — don't promote around a gap. 3. **Clear the reconcile queue.** Every merge conflict you resolve teaches the graph your naming reality. 4. **Curate the twenty facts that matter most.** The top tables' real definitions and owners deliver more agent-answer quality than a thousand auto-descriptions. 5. **Promote deliberately.** Authoritative facts are the ones agents will state without hedging — promote what you'd defend in a meeting. 6. **Expand scope** one schema at a time, repeating 2–5. ## How you know it's working Ask an agent a question whose answer lives in the graph — *"Who owns `fct_orders` and what does `net_revenue` actually mean here?"* A warm graph answers with provenance; a cold one says it doesn't know yet. Both are correct behavior — the difference is warmth, and the fix is the next phase of this page, not a workaround.