mlflow-mlops-migration

v2026.09.24

Guided workflow for taking any ML codebase — including one with no experiment tracking at all, or one full of MLflow 2-era idioms — to a production-grade open-source MLflow 3 setup with dev/staging/prod environments, registry-based promotion, and served models. Walks seven phases with a developer who may have zero MLflow 3 experience — assess the codebase (scripted read-only audit), model the registry domain (per-environment model names, aliases, gates), stand up tracking per environment, restructure training code to MLflow 3 idioms, wire evaluation-gated promotion, serve and smoke-test, then run the ongoing MLOps loop. Use when asked to set up MLflow, migrate to MLflow 3, productionize model training and serving, or design a dev/staging/prod MLOps cycle. Pairs with the sibling mlflow-3 rule pack for every API decision.

GitHub
安装命令
npx skhub add pproenca/mlflow-mlops-migration
Markdown
SKILL.md

MLflow MLOps Migration

A phased, gated workflow that turns an arbitrary ML codebase — however unstructured — into a production-grade open-source MLflow 3 setup covering the full MLOps cycle: tracked experiments, a domain-modelled registry, dev/staging/prod separation, evaluation-gated promotion, and served models. It is written to be driven with a developer who has no MLflow 3 experience: every phase produces a reviewable artifact before anything is changed, and every API decision defers to the sibling mlflow-3 rule pack (which is pinned to mlflow 3.15.1 and names the MLflow 2-era idioms this migration exists to remove).

When to Apply

Use this skill when:

  • A team wants MLflow (or has a messy/partial MLflow 2 setup) and needs the path to a production-grade MLflow 3 deployment — not just API fixes.
  • Training code exists but experiments are untracked, models are shipped by copying files, or "deployment" means a pickle in a bucket.
  • You are asked to design or review a dev/staging/prod model-promotion story.
  • An MLflow 2 → 3 migration touches infrastructure (stages, ./mlruns file stores, MLServer), not only client code.

Don't use it for a single API question — read the relevant mlflow-3 rule directly.

Workflow Overview

0 assess ─▶ 1 domain-model ─▶ 2 environments ─▶ 3 instrument ─▶ 4 promote ─▶ 5 serve ─▶ 6 operate
  audit        registry           tracking per      training code    eval-gated    validate,     retrain loop,
  report       naming, alias      env (dev local,   → MLflow 3       copy_model_   serve,        challenger,
  (script,     + gate design      stg/prod DB+S3    idioms (rule     version +     smoke-test    maintenance
  read-only)   (interview)        + auth)           pack)            alias flip    /invocations  (gated)
PhaseActionDeliverableRisk
0Run scripts/00-assess.sh <codebase> — read-only auditmlflow-assessment.md reportread-only
1Interview + domain modellingRegistry domain doc (names, aliases, gates)read-only
2Stand up tracking per environments; dev via scripts/scaffold-dev-tracking.shReachable tracking server(s), config.json filledwrite
3Restructure training code to MLflow 3 idioms (sibling rule pack)Refactored code, first LoggedModels registeredwrite
4Wire promotion — evaluate gate, tags, copy_model_version, alias flipPromotion script/CI jobwrite
5Serve — mlflow.models.predict, then serve/build-docker, smoke /invocationsServed model per environmentwrite
6Operate — retraining, challenger evaluation, maintenance (see workflow)Runbook habits, scheduled jobswrite
✓Run scripts/verify.sh after phases 2–5Pass/fail assertion reportread-only

Phases run in order — each has entry/exit criteria in references/workflow.md, and scripts/verify.sh is the exit gate for the infrastructure phases. Re-running any phase is safe: 00-assess.sh regenerates only its own report (and refuses to clobber anything else), scaffold-dev-tracking.sh refuses to overwrite (exit code 2 = already done), and verify.sh only reads. The one non-idempotent step is promotion's copy_model_version — see references/promotion.md for how to resume instead of re-copying.

Risk Level: Write

This workflow edits training code, writes infrastructure files, and stands up services. Guardrails:

  • Nothing in phase 0–1 modifies anything — always complete both before touching code or infra.
  • Confirm with the user before: starting/replacing any tracking server, rewriting a training entrypoint, flipping a prod @champion alias (dev/staging flips may be automated by the phase-4 pipeline), and exposing a serving endpoint beyond localhost.
  • Two maintenance commands are destructive and must be run only with explicit user confirmation and a stated reason: mlflow gc (permanently deletes soft-deleted runs and experiments — registry entities are untouched) and mlflow db upgrade (irreversible schema migration — snapshot the database first). A PreToolUse hook in hooks/hooks.json blocks both unless MLFLOW_MAINTENANCE_ACK=yes is set for that command, so they cannot run un-confirmed by accident.

Requirements

  • Python ≥ 3.10 with mlflow==3.15.1 installed in the project environment
  • bash, curl, jq — the scripts use them
  • uv — the serving phase uses --env-manager uv for fast isolated environment rebuilds (substitute virtualenv everywhere if uv is unavailable)
  • Docker + docker-compose — for the dev tracking stack and build-docker serving images
  • A database + object store per shared environment (staging/prod) — PostgreSQL/MySQL and S3/GCS/Azure; dev runs on the scaffolded local stack
  • The sibling mlflow-3 skill — phase 3 cites its rules; if it is not installed, read the MLflow 3 migration guide instead (the workflow still works, with more manual verification)

Setup

config.json starts empty. Phase 2 fills it (tracking URIs per environment, registry namespace, model name, serving URL). If fields are empty when a script needs them, the script says which ones — fill them via the _setup_instructions in the file.

Quick Reference

I need to…Go to
Audit what the codebase does todayscripts/00-assess.sh <dir> + references/assessment.md
Decide model names / aliases / gatesreferences/domain-modelling.md
Stand up dev tracking in one commandscripts/scaffold-dev-tracking.sh <dir>
Design staging/prod tracking topologyreferences/environments.md
Rewrite log_model / stages / evaluate callssibling mlflow-3 rules (log-*, reg-*, eval-*)
Build the promotion pipelinereferences/promotion.md
Serve and smoke-test a modelreferences/serving.md
Check the setup actually worksscripts/verify.sh
See every phase's entry/exit criteriareferences/workflow.md

Gotchas

See gotchas.md — failure points discovered while running this workflow, including the migrate-filestore SQLite-only target and the basic-auth bootstrap credentials.

Related Skills

  • mlflow-3 — the sibling library-reference rule pack this workflow cites at every API decision
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/.experimental/mlflow-mlops-migration

默认分支

master

最新提交

cf93c57

Tree SHA

afbb575