swarm-migrate — Cross-Repo Migration Swarm
One command, N repos, one coordinator, one ledger.
When to use
Use when the same transformation needs to land in 3 or more repos with the same shape (workflow bump, dependency upgrade, codemod, lint-rule introduction, secret rotation, runbook header). Don't use for one-repo work — that's implement. Don't use for novel exploration — that's brainstorm.
This skill exists because the 275-session insights showed 25 sessions burned coordinating PR cascades manually (M164 deploy-migration, M17 yg-mcp-core extraction, @v1 reusable workflow rollout across 14 repos). The pattern was always: pick a repo, branch, apply, push, watch CI, repeat. This automates the repeat.
vs CC
/workflows(2.1.154): CC's dynamic workflows orchestrate tens-to-hundreds of agents in the background and report via/workflows.swarm-migrateis different on purpose: it's a coordinator-led, foreground DAG with CI gates, conflict auto-rebase, and stop-on-novel-failure — you watch it and it stops on the first unexpected failure. Reach for CC/workflowswhen you want large-scale fire-and-forget background fan-out; reach forswarm-migratewhen each step needs a CI gate and a human-visible ledger. CC 2.1.202 adds a Dynamic workflow size setting in/config(small / medium / large agent counts) that advises/workflowson how many agents to spawn — it's a hint, not an enforced cap, so check it before hand-capping fan-out.
Inputs
A YAML spec at swarm-specs/<name>.yaml:
name: bump-actions-checkout-v4
description: "Pin @actions/checkout to v4 across all repos"
# Topology — repos in dependency order. Coordinator only proceeds
# to a downstream repo after every upstream parent has merged green.
repos:
- path: ~/coding/yonatan-hq/platform
upstream: []
- path: ~/coding/yonatan-hq/ventures/jobscraper
upstream: [platform] # waits for platform to merge first
# Transformation — applied identically per repo. The agent runs this
# inside the isolated worktree, then verifies with the next field.
transform:
type: codemod # codemod | regex | command
command: |
grep -rl 'actions/checkout@v3' .github/workflows | \
xargs sed -i '' 's|actions/checkout@v3|actions/checkout@v4|g'
# Verification — must pass before PR opens. Coordinator skips the repo
# if it fails locally (records skip-reason in ledger).
verify:
- command: "git diff --quiet"
expect: nonzero # must have changes
- command: "grep -r 'actions/checkout@v3' .github/workflows"
expect: nonzero # zero matches = clean
# PR shape — title, body, base branch
pr:
branch_prefix: chore/bump-checkout-v4
title: "chore(ci): pin @actions/checkout to v4"
body_file: swarm-specs/bump-actions-checkout-v4.pr.md
base: main
labels: [chore, ci]
# CI gate — coordinator waits for required checks to pass before
# moving downstream. Set to false for dry-run, or in repos without CI.
ci_gate:
required_checks: ["build", "test"]
timeout_minutes: 20
on_failure: pause # pause | skip | abort
# Limits
max_parallel: 4
abort_on_novel_failure: true
How it works
┌────────────────────────────────┐
│ COORDINATOR (you) │
│ reads spec → builds DAG → │
│ writes .swarm-state.json │
└─────────────────┬──────────────┘
┌────────────┼────────────┐
┌──────────┐ ┌──────────┐ ┌──────────┐
│ WORKER A │ │ WORKER B │ │ WORKER C │
│ (repo 1) │ │ (repo 2) │ │ (repo 3) │
└────┬─────┘ └────┬─────┘ └────┬─────┘
└─────────── isolated worktrees ─────┘
│ each: clone branch, transform,
│ verify, push, open PR, wait for CI, report
▼
┌────────────────────────────────┐
│ .swarm-state.json │
│ rolling ledger of {repo, │
│ status, pr_url, ci_state, │
│ last_action_at} │
└────────────────────────────────┘
Each worker is a Agent tool invocation (subagent type git-operations-engineer for plumbing or backend-system-architect for schema-flavored migrations). The coordinator (you, this skill) reads the ledger between waves and decides whether to release downstream waves or pause.
Push, do not poll (CC 2.1.224+)
Workers report state changes (CI green, CI red, novel failure, skipped) by MESSAGING
the coordinator via SendMessage instead of the coordinator re-reading the ledger
between waves:
Cross-session replies land in the parent (CC 2.1.248): when a subagent sends
SendMessageto another session, the reply is delivered to the parent session's conversation, never to the subagent; a subagent sends and moves on, the parent reads the answer. Cross-sessionSendMessage/ListAgentsalso work on Bedrock, Vertex and Foundry and with telemetry disabled (CC 2.1.248).
- Workers spawned as Agent-tool subagents use in-session
SendMessage(available since CC 2.1.77). - Workers running as separate sessions or
claude -pprocesses use cross-session messaging (CC 2.1.224, see chain-patterns Pattern 10). A-pworker must run withcrossSessionInbound: "accept"in--settingsto receive replies.
CRITICAL: cross-session delivery is not guaranteed - inbound controls can hold or
refuse a message. .swarm-state.json REMAINS the authoritative durable record, and
every state change is still written there. Messages are the low-latency signal; the
ledger is the truth.
Phase 1 — Spec validation
Load <spec-file.yaml>. Verify:
- Every
repos[].pathexists and is a git repo (usegit -C <path> rev-parsechecks). - The
transform.commandreturns 0 in a dry-run mode (ortransform.type: codemodresolves to a known codemod registered inswarm-specs/codemods/). - Every
upstreamreference points to a declared repo (no dangling deps). pr.body_fileexists and is non-empty.
If any check fails, abort and print the table of failures. Do NOT proceed.
Phase 2 — Topology sort
Build a DAG from upstream edges. Detect cycles → abort. Group nodes by topological wave (wave 0 = no deps, wave 1 = depends only on wave 0, …). Coordinator releases one wave at a time.
Write .swarm-state.json at the repo root running the skill:
{
"spec": "swarm-specs/bump-actions-checkout-v4.yaml",
"started_at": "2026-05-16T17:00:00Z",
"waves": [
{ "id": 0, "repos": ["platform"] },
{ "id": 1, "repos": ["jobscraper"] }
],
"repos": {
"platform": { "status": "pending", "pr_url": null, "ci_state": null, "last_action_at": null },
"jobscraper": { "status": "blocked", "blocked_on": ["platform"], "pr_url": null }
}
}
Phase 3 — Dispatch wave
For each repo in the current wave, in parallel (bounded by max_parallel):
- Worktree — create an isolated worktree at
<repo>/../<repo>-swarm-<spec-name>offorigin/<base>. Never mutate the live working tree. - Branch —
git checkout -b <branch_prefix>-<short-sha>. - Transform — run
transform.command(or apply codemod). Capture stdout to.swarm-logs/<repo>-transform.log. - Verify — run each
verify[].command, assert exit matchesexpect. On mismatch, mark reposkippedin ledger with reason, do not push. - Push + PR — push branch, open PR via
gh pr create. Update ledger with PR URL. - Watch CI — poll
gh pr checks <n>every 45s up toci_gate.timeout_minutes. Updateci_statein ledger on every state transition.
Use the Agent tool with subagent_type: ork:git-operations-engineer for steps 1–5 to keep main context lean. The coordinator only reads the ledger.
Phase 4 — Wave gate
After every wave, check the ledger:
- All
green→ release the next wave. - Any
pending CI→ keep polling. - Any
red CI→ consultci_gate.on_failure:pause→ halt the swarm, write a summary to.swarm-state.json, surface the failing logs, ask the user.skip→ mark repofailed-ci, continue with siblings (but block downstream unless they explicitly don't depend on this repo).abort→ terminate the swarm, leave open PRs as-is, never merge.
Phase 5 — Auto-rebase on conflicts
If a downstream repo's worker hits a merge conflict on rebase (because an upstream merged), the worker:
- Re-fetches the upstream's merge commit SHA.
- Attempts
git rebase origin/<base>. If clean → push, ledger update. - If conflicts → mark the conflict files in the ledger, do NOT auto-resolve, surface to the coordinator. Conflicts are the most common place auto-fixers ship broken code.
Phase 6 — Final report
When all waves complete (or the swarm pauses/aborts), emit a single markdown report under .swarm-logs/<spec-name>-report.md:
# Swarm report: bump-actions-checkout-v4
Completed: 12/14 repos · paused: 2 · duration: 47 min
| repo | status | PR | CI | duration |
|-----------------|---------|-------|--------|----------|
| platform | merged | #3456 | green | 8 min |
| jobscraper | merged | #281 | green | 6 min |
| ... |
| dormant-repo-1 | skipped | — | — | (no CI runner configured) |
| trading-ai | paused | #99 | red | (novel failure — see logs) |
## Novel failures (escalated)
- trading-ai #99: pyproject lockfile mismatch — see .swarm-logs/trading-ai-ci.log
Finish line. Done means: every repo in the spec has a PR open and green (never merged by the swarm), or is skipped with a reason, or paused with its novel failure logged, and the final report is written. Follow Read("../../shared/rules/long-run-protocol.md"): keep going when a step needs no input from the user, stop and ask only when you can't continue without them or before anything destructive, check each subagent's evidence before accepting it, and mark anything you couldn't confirm with where you looked.
Hard rules
- Never merge a PR. The swarm opens PRs; humans merge them. Auto-merge can be armed by the user with
gh pr merge --autopost-swarm if they want. - Never force-push. If a worker can't fast-forward, it pauses.
- Never roam outside the spec's declared
repos[]. Even if a transformation seems like it'd help elsewhere. - Always quarantine credentials. Workers run with the user's gh auth; the coordinator never logs tokens, just the URLs.
- Always respect existing branch protections. If
gh pr createfails because of required reviewers or other rules, that's a feature, not a bug to work around.
Failure modes you'll actually hit
| Mode | What it looks like | Mitigation |
|---|---|---|
| Stale lockfile | CI red on npm ci after dependency bump | Spec includes a post_transform.command: npm install step |
| Branch protection blocks PR creation | gh pr create exits non-zero | Coordinator marks repo blocked-by-protection, surfaces to user |
| Topology cycle | Phase 2 abort | Re-spec the upstream edges |
| Coordinator crash mid-flight | .swarm-state.json half-written | Skill is resumable: re-run with same spec, it reads the ledger and skips merged/green repos |
| Worker subagent hangs | No ledger update for >5 min | Coordinator times out the agent, marks repo worker-timeout, surfaces logs |
Related Skills
- Upstream —
brainstormto design the spec,visualize-planto ASCII-preview the DAG before dispatch. - Downstream —
verifyper repo after merge,/statusfor org-wide sweep,/ci-debugif a worker hits a CI red. - Composes with —
create-pr(each worker calls into it),github-operations(bulk-update labels/milestones post-swarm).
What this skill does NOT do
- Does not invent the spec. You write the spec; the skill executes it.
- Does not perform schema migrations across DBs (use a single-repo skill plus
database-patterns). - Does not orchestrate production deploys — open PRs only; deploy is a separate gate (the platform's deploy-operator).
- Does not bypass
create-pr's playground-gate rule — each PR body must include a playground reference if the repo enforces it.
Example invocation
# Dry-run: build the DAG, verify spec, do NOT push or open PRs
swarm-migrate swarm-specs/bump-actions-checkout-v4.yaml --dry-run
# Live: dispatch up to 4 workers in parallel
swarm-migrate swarm-specs/bump-actions-checkout-v4.yaml --max-parallel=4
# Resume after pause: same command, the ledger remembers
swarm-migrate swarm-specs/bump-actions-checkout-v4.yaml
Why this exists (one paragraph)
You ran 25 sessions in a single month coordinating cross-repo PRs by hand. The 14-repo @v1 workflow rollout, the M17 yg-mcp-core extraction, the M164 deploy-migration. Every one of those sessions had the same shape: a coordinator (you) holding the DAG in your head, dispatching workers (you, sequentially) in different terminal tabs, hand-rolling a status table in your notes. This skill makes the coordinator a YAML file and the workers parallel subagents. The DAG, the ledger, the auto-rebase, the wave gating — all the bookkeeping you were doing manually — get codified once. You write the spec, you walk away, you come back to a report.