adversarial-review

v2026.09.24

Adversarial second-pass review that tries to break code, designs, plans, or ADRs. Use when stakes are high and a normal review already ran.

GitHub
安装命令
npx skhub add laurigates/adversarial-review
Markdown
SKILL.md

Adversarial Review

A normal review asks "is this right?" and confirms intent. Adversarial review asks "how could this be wrong?" and hunts for the fault the first pass rationalised away. It is a second pass for residual risk on high-stakes work — not a replacement for the first pass, and not for low-stakes or reversible changes where it only manufactures busywork.

This is a thin posture, not a new checklist. It layers four moves — isolation, inverted objective, a triage gate, and a bounded loop — on top of the domain checklists that already live in code-review, security-audit, verify-before-plan, and cold-read-gate. It dispatches an isolated reviewer with the matching lens and then triages what comes back, because an agent told to find faults will invent them.

When to Use This Skill

Use this skill when...Use something else instead when...
Stakes are high and a normal review already passed, but residual risk remainsFirst-pass review of a diff/PR → code-quality-plugin:code-review
Red-teaming an architecture decision, ADR, or migration plan before commitVerifying plan facts (counts, paths, build state) → agent-patterns-plugin:verify-before-plan
Stress-testing a design's failure modes and invariantsPure security audit → agents-plugin security-audit / /security-review
Probing whether outward text survives a reader → wrap cold-read-gateLegibility of outward text on its own → agent-patterns-plugin:cold-read-gate
The change is hard to reverse and a missed fault is expensiveThe change is low-stakes, reversible, or throwaway — skip; it wastes tokens

The Four Moves

MoveWhat it meansBorrowed from
IsolationThe reviewer gets the artifact with no author context — bias can't leak incold-read-gate
Inverted objectiveBrief says "enumerate failure modes," never "confirm it works"this skill
Triage gateSeparate genuine faults from manufactured objections before actingcold-read-gate Step 3
Bounded loopOne revise round; a third means a structural problem the gate can't fixcold-read-gate Step 4

Model choice is the inverse of cold-read-gate. That skill uses haiku on purpose — the weak reader is the measurement instrument for legibility. Adversarial review wants opus: finding subtle faults is a reasoning task, not a low-context-reader simulation.

Parameters

Parse $ARGUMENTS:

  • Target (first positional) — what to attack: a path, a PR ref (#123 or URL), a file, or a free-text description of a plan/decision. If absent, default to the current diff (git diff + staged) and say so.
  • Focus directive (optional, free text after the target, e.g. focus on the failure path, ignore style) — biases the lens. It steers, never overrides: it must not cancel a live user boundary stated earlier in the session (the auto-mode.md conversation-boundary hazard).

Execution

Execute this adversarial review:

Step 1: Precondition gate

Confirm both hold before spending tokens:

  1. Stakes are high — the change is hard to reverse, or a missed fault is expensive. If not, stop and recommend a normal review instead.
  2. A first pass exists — a normal review/lint/test pass already ran. If not, run that first (code-review) — adversarial review is a second pass, and leading with it skips cheap, high-yield findings.

If either fails, say so and redirect rather than proceeding.

Step 2: Name the target and pick the lens

State in one line what is under review, then select the lens. The lens supplies the domain attack vocabulary — delegate to the owning skill's checklist rather than restating it:

Target typeLens / attack vocabularyDelegate the checklist to
Code, diff, PRLogic errors, edge cases, race conditions, error-swallowingcode-quality-plugin:code-review, code-review-checklist
Security surfaceInjection, authz gaps, secret exposure, trust boundariesagents-plugin security-audit, /security-review
Architecture / ADRCoupling, failure domains, blast radius, reversibilityblueprint adr-validate, adr-relationships
Plan / wave premisePremise truth, stale facts, name≠behaviouragent-patterns-plugin:verify-before-plan
Outward text / docsLegibility under zero contextagent-patterns-plugin:cold-read-gate
Research claimsUnverified or single-sourced claims.claude/skills/deep-research

Step 3: Dispatch the isolated reviewer

One Agent per lens. model: opus is the floor — see .claude/rules/agent-development.md § "Model Selection for Agents" — and model: fable is sanctioned here too: adversarial review is exactly the hardest delegated-reasoning work that guard accepts, so the reviewer must not be weaker than the author it is attacking. Dispatch lenses in parallel only when there are several and the session is not on a 1M-context model (every Fable 5.1 session, or Opus with the [1m] suffix) — the parallel-subagent rate-limit caveat in skill-fork-context.md. Template:

subagent_type: general-purpose
model: opus   # floor; fable also sanctioned — see agent-development.md
prompt: |
  You are an adversarial reviewer. Your ONLY objective is to find ways this
  is WRONG, fragile, or unsafe — do not confirm that it works, do not
  praise it. Assume a fault exists and locate it.

  Read ONLY the target (no scope beyond it):
  <target path / diff / pasted plan>

  Attack along this lens: <lens from Step 2, with its checklist>.
  <focus directive, if any>

  Produce:
  1. FAULTS — each as: SEVERITY=critical|high|medium  EVIDENCE=<file:line or
     quoted claim>  FAILURE=<the concrete way it breaks>.
  2. ASSUMPTIONS-ATTACKED — load-bearing assumptions you tried to falsify,
     and whether each held.
  3. VERDICT: exactly one of `sound` | `flawed`.
  Cite evidence for every fault. Your final message is the deliverable.

Step 4: Triage — genuine fault vs manufactured objection

The reviewer was told to find faults, so some "faults" are noise. Triage before acting (this is the load-bearing step — skipping it is how adversarial review sends you down the wrong path):

Genuine fault — act on itManufactured objection — drop it
A concrete input/state that breaks the code, with evidenceA defense against an input the contract makes impossible
A failure mode with real blast radiusA hypothetical with no realistic trigger
A violated invariant or unhandled error pathStyle/preference dressed up as a fault
A premise that is actually falseRe-litigating a trade-off already decided with rationale
A missing edge case the spec impliesScope the change didn't touch and isn't responsible for

Step 5: Report and bound the loop

Emit a prioritised report: target, verdict, surviving genuine faults (severity-ordered, with evidence and a suggested fix), and explicitly note the objections you dropped in triage and why. Apply or hand off the genuine fixes. Re-dispatch a fresh reviewer only if the verdict was flawed; do not loop more than twice — a third round means a structural problem the review can't resolve.

Anti-patterns

MistakeCorrect approach
Running it as a first passIt's a second pass — a normal review runs first (Step 1)
Using it on low-stakes / reversible workSkip it; the precondition gate exists to say no
Acting on every objection the reviewer raisesTriage first (Step 4); the inverted objective guarantees noise
Using a weak model "to be tougher"Opus — subtle faults are a reasoning task (inverse of cold-read-gate)
Restating each lens's checklist inlineDelegate to the owning skill; this is a posture, not a checklist
Looping until the reviewer goes silentOne revise round; persistent faults = structural problem

Related

  • REFERENCE.md — audit sweeps: find, verify adversarially, then resolve settled facts before editing; apply an undecided finding only where the facts make it policy-neutral, never "TBD" or a plausible owner
  • cold-read-gate — the isolation + triage + bounded-loop pattern this skill generalises (legibility lens; uses haiku)
  • verify-before-plan — adversarial review of premises; the plan/wave lens delegates here
  • execution-grounded-review — the sibling verifier for running behaviour against acceptance criteria (this skill attacks a design; that one grounds each criterion in execution evidence)
  • code-quality-plugin:code-review — the first-pass review this layers on top of
  • agents-plugin security-audit / /security-review — the security lens
  • .claude/rules/terminology.md — defines Adversarial review and Red-team
  • .claude/rules/skill-fork-context.md — the [1m] parallel-dispatch caveat
  • .claude/rules/loop-integrity.md — looping skills delegate their stop-condition judgement to an isolated reviewer like this one (Pillar 1)
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

agent-patterns-plugin/skills/adversarial-review

默认分支

main

最新提交

1668324

Tree SHA

b2d4cc3