MLOps & Models (dw-ml)
The MLOps & Models agent manages the machine-learning lifecycle end to end: experiment tracking, a versioned model registry with stage management, feature pipelines with drift statistics, model explainability, and guarded rollout with A/B testing. It treats a model as one service in a graph — data, features, training, registry, serving, monitoring — not a notebook artifact, and it treats a feature as a versioned product with an entity key, null semantics, and a freshness expectation, not a column name in a file.
Its engineering instincts are the ones that survive production: assume train–serve skew until proven otherwise, audit for leakage before believing any offline gain, require a reproducible run before promotion, and prove the rollback path before any traffic shifts. A “better model” with no kill switch is, in this agent’s view, not shippable.
Key capabilities
Section titled “Key capabilities”- Experiment tracking.
create_experiment,log_metrics, andcompare_experimentsgive you MLflow-compatible run tracking with step-based training curves and cross-run comparison including significance tests. - Model registry with stages.
register_modelandget_model_versionsversion models through development, staging, and production, with stage history and lineage. - Budget-aware AutoML training.
ml_automl_trainis the real-compute path: you set a time budget, not an algorithm, and a FLAML-based search returns the best estimator, tuned hyperparameters, and real cross-validation scores, persisting a versioned model artifact. - Feature pipelines and stats.
create_feature_pipelinegenerates scheduled feature jobs (sinks include Snowflake, BigQuery, Redis);get_feature_statsreports distributions, null rates, and drift scores against training baselines. - Explainability.
explain_modelproduces SHAP-style feature-importance reports (permutation and gain methods also available), down to individual-prediction explanations. Today these are deterministic simulations for evaluating the workflow — not values computed against your live model. - Drift detection.
detect_model_driftseparates data drift from concept drift using KS, PSI, and Chi-squared tests with configurable thresholds. - Guarded rollout and A/B testing.
deploy_modelsupports canary, shadow, and blue-green strategies;ab_test_modelsconfigures traffic splits and reports metric comparison with statistical significance. - Feature suggestions and model selection.
suggest_featuresranks feature-engineering candidates from a dataset profile;select_modelrecommends architectures for the problem type and constraints.
Example prompts
Section titled “Example prompts”“Create an experiment for the churn model and run an AutoML search with a 20-minute budget.”
“Compare the last three experiment runs — which wins on calibration and the worst slice, not just AUC?”
“Set up a feature pipeline for these five features with a daily refresh into Snowflake.”
“Has the fraud model drifted in the last 7 days? Separate data drift from concept drift.”
“A/B test v12 against the champion at a 10% split, primary metric precision@k.”
Connect it to your stack
Section titled “Connect it to your stack”- Warehouses and feature sinks — Snowflake, BigQuery, Databricks for training data and feature storage.
- ML tooling — MLflow and Weights & Biases are in the ML category of the connector catalog.
- Catalog — DataHub and dbt supply the feature and metric definitions the agent resolves instead of re-deriving.
Works before you connect anything
Section titled “Works before you connect anything”The agent starts in 🟡 Evaluation on built-in sample data — you can walk the full lifecycle, from experiment to registry to a simulated rollout, before any credential exists. It earns 🟢 Connected per system through a passing live test. See Verify your setup.
Limits, honestly
Section titled “Limits, honestly”train_modelis a deterministic simulation kept for compatibility — for a real fit, use the AutoML path (ml_automl_train). That path requires a Python runtime with FLAML available; without one it fails with a clear error rather than faking metrics, and a timed-out search is reported as a failed run, not a shippable best-so-far.- Training budgets are time and trials — never dollars.
- In 🟡 Evaluation, registry, deployment, and monitoring act on the sample environment; nothing reaches real serving infrastructure until connections are verified.
- The agent’s own discipline is a limit you’ll feel: no promotion without a reproducible run, and no rollout without a tested rollback path. It will refuse shortcuts rather than register an unreproducible champion.
- Data-quality breaks it finds are handed to the Quality agent rather than silently worked around.