mlflow-mlops-migration

v2026.09.24

Guided workflow for taking any ML codebase — including one with no experiment tracking at all, or one full of MLflow 2-era idioms — to a production-grade open-source MLflow 3 setup with dev/staging/prod environments, registry-based promotion, and served models. Walks seven phases with a developer who may have zero MLflow 3 experience — assess the codebase (scripted read-only audit), model the registry domain (per-environment model names, aliases, gates), stand up tracking per environment, restructure training code to MLflow 3 idioms, wire evaluation-gated promotion, serve and smoke-test, then run the ongoing MLOps loop. Use when asked to set up MLflow, migrate to MLflow 3, productionize model training and serving, or design a dev/staging/prod MLOps cycle. Pairs with the sibling mlflow-3 rule pack for every API decision.

GitHub
Install command
npx skhub add pproenca/mlflow-mlops-migration
Markdown
SKILL.md

MLflow MLOps Migration

A phased, gated workflow that turns an arbitrary ML codebase — however unstructured — into a production-grade open-source MLflow 3 setup covering the full MLOps cycle: tracked experiments, a domain-modelled registry, dev/staging/prod separation, evaluation-gated promotion, and served models. It is written to be driven with a developer who has no MLflow 3 experience: every phase produces a reviewable artifact before anything is changed, and every API decision defers to the sibling mlflow-3 rule pack (which is pinned to mlflow 3.15.1 and names the MLflow 2-era idioms this migration exists to remove).

When to Apply

Use this skill when:

  • A team wants MLflow (or has a messy/partial MLflow 2 setup) and needs the path to a production-grade MLflow 3 deployment — not just API fixes.
  • Training code exists but experiments are untracked, models are shipped by copying files, or "deployment" means a pickle in a bucket.
  • You are asked to design or review a dev/staging/prod model-promotion story.
  • An MLflow 2 → 3 migration touches infrastructure (stages, ./mlruns file stores, MLServer), not only client code.

Don't use it for a single API question — read the relevant mlflow-3 rule directly.

Workflow Overview

0 assess ─▶ 1 domain-model ─▶ 2 environments ─▶ 3 instrument ─▶ 4 promote ─▶ 5 serve ─▶ 6 operate
  audit        registry           tracking per      training code    eval-gated    validate,     retrain loop,
  report       naming, alias      env (dev local,   → MLflow 3       copy_model_   serve,        challenger,
  (script,     + gate design      stg/prod DB+S3    idioms (rule     version +     smoke-test    maintenance
  read-only)   (interview)        + auth)           pack)            alias flip    /invocations  (gated)
PhaseActionDeliverableRisk
0Run scripts/00-assess.sh <codebase> — read-only auditmlflow-assessment.md reportread-only
1Interview + domain modellingRegistry domain doc (names, aliases, gates)read-only
2Stand up tracking per environments; dev via scripts/scaffold-dev-tracking.shReachable tracking server(s), config.json filledwrite
3Restructure training code to MLflow 3 idioms (sibling rule pack)Refactored code, first LoggedModels registeredwrite
4Wire promotion — evaluate gate, tags, copy_model_version, alias flipPromotion script/CI jobwrite
5Serve — mlflow.models.predict, then serve/build-docker, smoke /invocationsServed model per environmentwrite
6Operate — retraining, challenger evaluation, maintenance (see workflow)Runbook habits, scheduled jobswrite
✓Run scripts/verify.sh after phases 2–5Pass/fail assertion reportread-only

Phases run in order — each has entry/exit criteria in references/workflow.md, and scripts/verify.sh is the exit gate for the infrastructure phases. Re-running any phase is safe: 00-assess.sh regenerates only its own report (and refuses to clobber anything else), scaffold-dev-tracking.sh refuses to overwrite (exit code 2 = already done), and verify.sh only reads. The one non-idempotent step is promotion's copy_model_version — see references/promotion.md for how to resume instead of re-copying.

Risk Level: Write

This workflow edits training code, writes infrastructure files, and stands up services. Guardrails:

  • Nothing in phase 0–1 modifies anything — always complete both before touching code or infra.
  • Confirm with the user before: starting/replacing any tracking server, rewriting a training entrypoint, flipping a prod @champion alias (dev/staging flips may be automated by the phase-4 pipeline), and exposing a serving endpoint beyond localhost.
  • Two maintenance commands are destructive and must be run only with explicit user confirmation and a stated reason: mlflow gc (permanently deletes soft-deleted runs and experiments — registry entities are untouched) and mlflow db upgrade (irreversible schema migration — snapshot the database first). A PreToolUse hook in hooks/hooks.json blocks both unless MLFLOW_MAINTENANCE_ACK=yes is set for that command, so they cannot run un-confirmed by accident.

Requirements

  • Python ≥ 3.10 with mlflow==3.15.1 installed in the project environment
  • bash, curl, jq — the scripts use them
  • uv — the serving phase uses --env-manager uv for fast isolated environment rebuilds (substitute virtualenv everywhere if uv is unavailable)
  • Docker + docker-compose — for the dev tracking stack and build-docker serving images
  • A database + object store per shared environment (staging/prod) — PostgreSQL/MySQL and S3/GCS/Azure; dev runs on the scaffolded local stack
  • The sibling mlflow-3 skill — phase 3 cites its rules; if it is not installed, read the MLflow 3 migration guide instead (the workflow still works, with more manual verification)

Setup

config.json starts empty. Phase 2 fills it (tracking URIs per environment, registry namespace, model name, serving URL). If fields are empty when a script needs them, the script says which ones — fill them via the _setup_instructions in the file.

Quick Reference

I need to…Go to
Audit what the codebase does todayscripts/00-assess.sh <dir> + references/assessment.md
Decide model names / aliases / gatesreferences/domain-modelling.md
Stand up dev tracking in one commandscripts/scaffold-dev-tracking.sh <dir>
Design staging/prod tracking topologyreferences/environments.md
Rewrite log_model / stages / evaluate callssibling mlflow-3 rules (log-*, reg-*, eval-*)
Build the promotion pipelinereferences/promotion.md
Serve and smoke-test a modelreferences/serving.md
Check the setup actually worksscripts/verify.sh
See every phase's entry/exit criteriareferences/workflow.md

Gotchas

See gotchas.md — failure points discovered while running this workflow, including the migrate-filestore SQLite-only target and the basic-auth bootstrap credentials.

Related Skills

  • mlflow-3 — the sibling library-reference rule pack this workflow cites at every API decision
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/.experimental/mlflow-mlops-migration

Default branch

master

Latest commit

cf93c57

Tree SHA

afbb575