execution-grounded-review

v2026.09.24

Execution-grounded review: run tests first, trace each acceptance criterion to execution evidence. Use when verifying an implementation meets spec.

GitHub
安装命令
npx skhub add laurigates/execution-grounded-review
Markdown
SKILL.md

Execution-Grounded Review

A normal review reads the diff and asks "does this look right?" — and an implementation can look complete while a criterion silently fails. This skill refuses to grade an implementation on appearance: it runs the suite first, then traces each acceptance criterion to execution evidence — a test that actually exercised it, or observed behaviour — and marks anything it cannot back with execution as UNVERIFIED rather than passing it on faith.

It is the execution-grounded verifier in the agent-patterns review family: where adversarial-review attacks a design for faults and cold-read-gate measures whether text survives a reader, this skill verifies that running code meets its stated acceptance criteria. It is the reusable independent verifier a judgement-based loop gate delegates to (.claude/rules/loop-integrity.md, Pillar 1).

When to Use This Skill

Use this skill when...Use something else instead when...
Verifying an implementation meets explicit acceptance criteria, proven by running itFirst-pass review of a diff → code-quality-plugin:code-review
Gating a loop/phase done on an independent check of behaviourRed-teaming a design or ADR for faults → agent-patterns-plugin:adversarial-review
Confirming a fix actually fixes the reported failure (not just compiles)Checking premises/facts before work starts → agent-patterns-plugin:verify-before-plan
Closing the loop on "is every requirement actually covered by a test?"Legibility of outward text → agent-patterns-plugin:cold-read-gate

The Stance

MoveWhat it means
Execute firstRun the full suite + typecheck + lint before any verdict. A criterion is PASS only with execution evidence — never "the code looks like it does this".
Trace each criterionOne ledger row per acceptance criterion: premise → evidence (file:line / test name / observed output) → verdict.
No silent passA criterion with no execution backing is UNVERIFIED (a coverage gap to surface), not an assumed pass.
Match the production sequenceFor a round-trip / determinism / reproducibility / idempotence claim, a passing test is evidence only if its operation sequence reproduces the real production call path — not a convenient shorter one (see Step 3a).
Intent-starved verifierThe isolated verifier reads the criteria, the diff, and the captured execution evidence — not the author's plan narrative or rationale, which would let it rationalise a pass.
Bounded loopOne revise round on fail; a third means a structural problem the gate can't resolve.

Model is opus — like adversarial-review, the inverse of cold-read-gate: building an accurate requirement→evidence ledger is a reasoning task.

Parameters

Parse $ARGUMENTS:

  • Target (first positional) — what to verify: a path, a PR ref (#123 or URL), explicit files, or absent. If absent, default to the current change (git diff HEAD + staged) and say so.
  • --criteria <file> (optional) — a file of acceptance criteria. If absent, gather criteria from the task/plan in context (the acceptance criteria stated for this change) and echo them back before verifying, so the user can correct the list the skill is grading against.

Execution

Execute this execution-grounded verification:

Step 1: Run the suite first

Before reading the diff for "correctness", establish ground truth by execution. Detect and run the project's full suite + typecheck + lint, capturing output to a scratch file (this is the execution evidence the verifier grades against):

StackSuiteTypecheckLint
Node/TSnpm test (or npx vitest run)npx tsc --noEmitnpx biome check
Pythonuv run pytest -quv run ty checkuv run ruff check
Rustcargo testcargo checkcargo clippy
Gogo test ./...go vet ./...—

Record the exit codes and failing-test names. A red suite is itself an independent signal — a failing test does not care how hard the author worked (.claude/rules/loop-integrity.md).

Step 2: Name the criteria

State, in one numbered list, the acceptance criteria under verification (from --criteria or context). If the list is empty, stop and ask for it — there is nothing to ground a verdict in. Each criterion is one ledger row in Step 3.

Step 3: Dispatch the intent-starved verifier

One Agent, model: opus, reading only the criteria, the diff, and the captured execution-evidence file — not the author's reasoning. Bind its output to the LEDGER schema below and paste that schema verbatim into the brief: a schema forces a determinate answer on every row where prose lets a row go quietly unanswered.

The LEDGER schema

{
  "type": "object",
  "required": ["rows", "coverage", "verdict"],
  "properties": {
    "rows": {
      "type": "array",
      "items": {
        "type": "object",
        "required": ["criterion", "evidence", "evidenceSpan", "verdict", "sequenceMatchesProduction"],
        "properties": {
          "criterion": { "type": "string" },
          "evidence": { "type": "string" },
          "evidenceSpan": { "type": "string" },
          "verdict": { "enum": ["PASS", "FAIL", "PARTIAL", "UNVERIFIED"] },
          "sequenceMatchesProduction": { "enum": ["yes", "no", "not-applicable"] }
        }
      }
    },
    "coverage": { "type": "string" },
    "verdict": { "enum": ["pass", "fail"] }
  }
}

evidence is the test name / file:line / observed output drawn from the execution-evidence file, or the literal none. evidenceSpan is where that evidence sits in the execution-evidence file — <file>:<startLine>-<endLine> — or the literal none; it lets the caller check an attribution instead of trusting it (Step 3b). It carries no maxLength or pattern on purpose: a locator rejected by a length cap gets resent, not fixed (#2280). coverage is <#rows with PASS/FAIL evidence> / <total rows>.

Row verdictMeaning
PASSexecution evidence demonstrates the criterion holds
FAILexecution evidence demonstrates it is violated — name the concrete failing input/test
PARTIALcovered for some inputs; a stated edge case is unhandled
UNVERIFIEDno execution exercises this criterion (a coverage gap) — never pass a row because the code "looks right"

sequenceMatchesProduction is required on every row, and that requirement is the entire gain of the schema. Step 3a's check is the one a verifier skips silently when it is only prose, because a green test looks like evidence — a required enum makes "I did not check" unrepresentable.

ValueWhen
yesthe test's operation sequence reproduces the real production call path
nothe test passes over a shorter or rearranged sequence than production uses → the row's verdict becomes UNVERIFIED
not-applicablethe criterion makes no round-trip / determinism / reproducibility / idempotence claim

The overall verdict is pass only when every row is PASS and no row is sequenceMatchesProduction: "no". Any FAIL, any UNVERIFIED, or any sequence divergence makes it fail.

Template:

subagent_type: general-purpose
model: opus
prompt: |
  You verify whether an implementation meets its acceptance criteria, grounded
  in EXECUTION EVIDENCE. Read ONLY these inputs (no other files, no repo
  exploration beyond resolving evidence cited below):
    - Acceptance criteria: <numbered list from Step 2>
    - The change under review: <diff / file paths>
    - Execution evidence (suite/typecheck/lint output): <scratch file path>

  Do NOT read the author's plan, commit narrative, or rationale — grade the
  behaviour, not the intent.

  Emit ONE object conforming to this schema, with one row per criterion:
    <paste the LEDGER schema verbatim>

  For every row, decide `sequenceMatchesProduction` explicitly: identify what
  real call sequence exercises the claim in production and confirm the test
  reproduces that sequence, not just a convenient shorter one.

  When a criterion fails and the evidence file is long, locate before you read:
  `grep -n` it for anchors from the criterion (test name, assertion text, error
  class, changed file path), read only a window around each hit, and let a
  candidate that does not explain the failure narrow the next query. Stop after
  3 search rounds or 5 windows for that criterion; if nothing explains it, set
  its `evidenceSpan` to "none" and say the cause was not located.
  Cite evidence and its `evidenceSpan` for every row. Your final message is the
  deliverable.

No workflow harness here — deliberately. The schema is the whole delta; agent count stays at one. This skill is the Pillar-1 oracle other loops delegate their stop condition to, so a fan-out design would multiply through every iteration of every loop in the repo — the one place where per-invocation cost compounds rather than adds (.claude/rules/workflow-vs-skill.md).

For several independent targets, dispatch one verifier per target in a single-message parallel Agent batch — except on a 1M-context model (every Fable 5.1 session, or Opus with the [1m] suffix), where the concurrent- subagent rate-limit caveat applies (skill-fork-context.md); run those sequentially. Do not set context: fork — the caller needs the ledger in the main context to act on it.

Step 3a: Match the test's operation sequence to production

"A passing test exists for this claim" is not sufficient evidence for a round-trip, determinism, reproducibility, or idempotence claim. The test can pass while the design is broken, because a narrower hand-built repro silently avoids the exact ordering that would expose a divergence. The tell is when the test's sequence of operations differs from the real call path production uses.

So for any such claim, add a step to the verifier's brief:

Identify what real call sequence exercises this claim in production, and confirm the test under review reproduces that sequence — not just a convenient shorter one.

Stateful / RNG-dependent code is the high-risk class — lazy initialization, global mutable RNG state, and caching all defer or share observable state, so when an operation runs relative to its neighbours changes the result. A test of the shape construct → forward immediately and a production path of construct → generate batches (consuming lazy RNG) → first forward draw their deferred state at the same relative point in each stream, so the round-trip "works" with no error — while a real trained model's reload diverges. The passing test is real; it just exercises the one ordering that can't see the bug.

Grade such a claim UNVERIFIED until the test reproduces the production sequence, even though a green test exists — that is the row whose sequenceMatchesProduction is no.

Step 3b: Attribute a failure by searching the trace, not reading it

When a criterion fails on a long run, its evidence is usually sparse and spread across actions far apart in the trace. Reading end to end gets worse as the trace grows: the relevant lines are diluted by irrelevant ones, and the reader settles on the most recent plausible cause rather than the actual one. Summarising first does not help — a summary is exactly where sparse evidence is dropped. So when a row is heading for FAIL or PARTIAL and the execution-evidence file is longer than one pass (roughly 300+ lines), the verifier searches instead:

  1. Locate first. grep -n the evidence file for anchors taken from the criterion — the test name, the assertion text, the error class, the changed file's path. Each hit is a candidate.
  2. Read narrowly. Open only a window around each candidate (about ±20 lines), never the file end to end.
  3. Narrow, don't stop. A candidate that does not explain the failure feeds the next query — drop its anchor, add the symbol it pointed at — rather than ending the search on the nearest plausible line.
  4. Report the span. The window that explains the failure becomes the row's evidenceSpan, so the caller can open that range and check the attribution.

Attribution bound: at most 3 search rounds and 5 candidate windows per criterion — the hard ceiling .claude/rules/loop-integrity.md ("Bounding runaway") requires of every loop. When the bound is reached with no explaining span, the row records evidenceSpan: "none" and states that the cause was not located within the bound. That row goes to a human; it is never filled with the most plausible-looking line. A named failing test in evidence still stands — the missing span says only that its cause was not pinned down.

Step 4: Triage against over-correction

The verifier grades strictly, so guard both failure modes before acting — neither talk yourself into passing broken code, nor into failing correct code:

Act on itDrop it
A FAIL with a named failing test/inputA FAIL on a requirement the spec never stated
An UNVERIFIED criterion → write/run the missing test, then re-gradeAn UNVERIFIED on behaviour outside the change's responsibility
A PARTIAL where a stated edge case is unhandledStyle/preference dressed up as a criterion failure
A coverage gap on a load-bearing criterionA hypothetical input the contract makes impossible
A round-trip/determinism test whose sequence diverges from production (Step 3a)A sequence difference that provably can't affect the claim's outcome

Step 5: Report and bound the loop

Emit the ledger: target, per-criterion rows with evidence and evidenceSpan, COVERAGE, and the overall verdict. Apply or hand off the genuine fixes (closing UNVERIFIED rows by adding the missing test counts as a fix). Re-run from Step 1 only if the verdict was fail; do not loop more than twice — a third round means a structural problem the gate can't resolve, which is the signal to surface to a human, not to keep grinding.

Anti-patterns

MistakeCorrect approach
Grading the diff without running anythingExecute first (Step 1) — appearance is not evidence
Passing a criterion because the code "looks like it does that"No execution evidence → UNVERIFIED, not pass
Passing a round-trip/determinism claim because "a test exists and passes"Confirm the test's operation sequence matches the production call path (Step 3a)
Reading a long trace end to end and naming the latest plausible causeLocate with grep -n, read narrow windows, report the span — within the attribution bound (Step 3b)
Feeding the verifier the author's plan/rationaleIntent-starved inputs — criteria + diff + execution evidence only
Inventing requirements the spec never statedTriage (Step 4) — FAIL only on listed criteria
Looping until the verifier goes quietOne revise round; persistent fail = structural problem

Related

  • adversarial-review — attacks a design for faults; this skill verifies running behaviour against criteria
  • verify-before-plan — verifies premises before work; this verifies outcomes after
  • cold-read-gate — the isolation + triage + bounded-loop pattern this skill reuses (legibility lens; uses haiku)
  • code-quality-plugin:code-review — the first-pass review this layers on top of
  • workflow-orchestration-plugin:workflow-checkpoint-refactor — a loop whose phase gate delegates its independent verdict here
  • .claude/rules/loop-integrity.md — Pillar 1: a loop's stop condition is judged by an independent verifier like this one, not the worker. Keep this literal path in the body: scripts/check-loop-integrity.sh requires the token loop-integrity.md in this file, so a later "tighten the Related section" edit that drops it fails the build.
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

agent-patterns-plugin/skills/execution-grounded-review

默认分支

main

最新提交

1668324

Tree SHA

b2d4cc3