ci-forensics

v2026.09.25

This skill should be used when reading CI logs and diagnosing failures — fetching Actions logs through the GitHub MCP, classifying failures by cause, and distinguishing flaky tests from real regressions.

GitHub
Install command
npx skhub add thelobbi/ci-forensics
Markdown
SKILL.md

CI Forensics

Fetching logs — the correct sequence

1. mcp__github__actions_list   method: "list_workflow_runs", branch: <head ref>
2. mcp__github__get_job_logs   run_id: <run>, failed_only: true
3. mcp__github__get_job_logs   job_id: <job>, return_content: true, tail_lines: 200

Known pitfalls in these tools:

  • actions_list and actions_get are multiplexed behind a method discriminator. Calling either without method errors.
  • actions_get also needs resource_id (not run_id) and returns metadata, not logs. It is the wrong tool for log triage.
  • get_check_run fetches one check run by checkRunId; it has no list mode. For "is CI green on this sha", use actions_list → list_workflow_runs and filter by head_sha.
  • list_workflow_runs can exceed the token limit and be spilled to a file. Parse that file with node -e, do not read it inline.

Failure taxonomy

Classify before changing anything. Fixing an infra flake by editing source code produces a red PR with extra churn.

ClassSignalsAction
Real regressionTraces to a file/line in the diff; reproduces locally; the test is new or newly touchedFix + regression test
FlakySame commit sha, different result; timing/ordering/port/network in the traceConfirm across ≥3 runs, then quarantine
InfraRunner OOM, disk full, image pull failure, registry 5xx, action download timeoutRe-run. Escalate after 2. Never edit code
Base brokenSame job red on the base at a commit predating the branch pointReport once, wait for recovery
Config driftFailure in workflow YAML, action inputs, or a missing secretFix the workflow
DependencyAppeared with no source change; a transitive package publishedPin or roll back

Establishing "base broken"

Do not infer it from error text. Check the base branch's own runs for the same job at a commit predating this branch's point of divergence. Red there too → base breakage. Otherwise it is yours.

Flaky vs real

EvidenceStrength
Same commit sha, different resultsConclusive
Passed on a neighbouring run, no relevant changeStrong
Fails only under parallel executionStrong — ordering or shared state
Fails only on one OS or runtime versionMedium — environment dependence
"It passed when I re-ran it"Weak — never sufficient alone

A test that has never passed is not flaky. It is broken.

Failure rate hints at mechanism: ~50% is usually an async race; ~2% is usually resource pressure or a real-time dependency.

Sources of nondeterminism

Real time · unseeded randomness · test ordering and shared module state · fixed ports · shared database rows · reused temp paths · missing awaits · real network calls · resource pressure.

Never

  • Never disable, skip, or .only a test to make CI pass. .only silently disables the rest of the file.
  • Never re-run a job to make a failure disappear without recording why it failed.
  • Never claim CI is fixed without a green run to point at.

See also

  • actions-authoring — fixing the workflow itself
  • ../commands/ci.md — the drive-to-green loop
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.25

Published

Sep 25, 2026

Category

Uncategorized

License

MIT

Source path

plugins/delivery-orchestrator/skills/ci-forensics

Default branch

main

Latest commit

2f1269c

Tree SHA

629e050