ai-research-reproduction
Purpose
Guide README-first deep learning reproduction toward the smallest trustworthy run with auditable evidence. Preserve documented meaning; record assumptions, deviations and blockers instead of changing semantics to manufacture success. Load specialized references only for a concrete uncertainty.
Fast Path
For a routine bounded run, keep the control path short:
- Read the target README and only the target test/config/source needed to understand the documented command.
- Run
scripts/orchestrate_repro.py --repo <repo> --plan-only --agent-outputwith any explicit user timeout bound (--timeoutor--train-timeout) already supplied; reviewcommand_candidates, the selectedcmd-XX, side-effect contract, selection fingerprint, and returnedreviewed_run_args. With no--output-dir, later evidence goes to<repo>/repro_outputsregardless of caller cwd. - Run the selected candidate, or another reviewed candidate, with
--run-selected --command-id <cmd-XX> --plan-fingerprint <fingerprint> --agent-outputplus requested timeout/metric/source-adjacent options. Preserve an explicit user command-timeout bound instead of silently making it stricter on a routine trusted run.--timeoutlimits the target command; do not wrap the whole orchestrator in an equal or shorter external timeout, because it still needs time to terminate children and write terminal evidence. A changed command set fails closed; setup/download commands are never target candidates. - Run
--verify-output --agent-output; inspect detailed evidence files only when verification fails or the result is partial/blocked. - Deliver the bounded result and stop.
For a host with short tool-call deadlines, rerun planning with --include-agent-handoff and follow references/agent-job.md; otherwise keep the direct path above. Job completion is not task acceptance, and uncertain state is never a reason for automatic replay.
Do not inspect orchestrate_repro.py, annotate_readme.py, _bundled/, writers, or runtime internals on a normal success path. Inspect implementation only for a concrete blocker, unexpected side effect, bundle-integrity failure, or unresolved safety question. Use scripts/doctor.py for first-use environment/install diagnostics. Executed commands keep full lifecycle/log evidence under repro_outputs/_runtime/<run_id>/.
Fit
Use this skill for repository-grounded, multi-phase trusted reproduction where the goal is a small reproducible target. Do not use it for paper summaries, generic setup, isolated scanning, standalone commands, open-ended research design, or explicitly authorized candidate exploration.
Trusted Target Selection
Choose the smallest target that can honestly demonstrate repository-grounded reproduction:
- documented inference
- documented evaluation
- documented training startup or partial verification
- full training only after explicit user confirmation
Treat README guidance as the primary reproduction intent. Use repository files
to clarify the README, not to silently replace it. When the README and paper
conflict, record the conflict and use paper-context-resolver only for the
narrow reproduction-critical gap.
Workflow
- Treat README guidance as primary; extract and select the minimum trustworthy target.
- Use setup/assets only for target-specific prerequisites and
analyze-projectonly when structural clarification is needed. - Use
minimal-run-and-auditfor inference/evaluation/smoke andrun-trainfor training startup, kickoff, or resume; direct execution is the default. - Pause before fuller training or changes to dataset, split, checkpoint, preprocessing, metric, loss, model semantics, or interpretation.
- Award
result-matchonly against explicit expected metrics and tolerance; process success alone is not reproduction success. - Write the evidence bundle, return the requested bounded result, and stop; optional stages are not automatic follow-up work.
Patch Boundary
Prefer no repository edits. If edits are needed, keep them conservative and auditable:
- Try command-line arguments, environment variables, path fixes, dependency version fixes, or dependency-file fixes before code changes.
- Reproduction fixes are allowed when needed, but they must not be hidden. State what changed, why it was necessary, whether it changes scientific meaning, and whether it affects comparability with the paper, README, or baseline.
- Avoid changing model architecture, core inference semantics, training logic, loss functions, or experiment meaning.
- If repository files must change, create a branch named
repro/YYYY-MM-DD-short-task, keep verified patch commits sparse, and record README-fidelity impact inPATCHES.md.
See references/patch-policy.md.
Outputs
Always target repro_outputs/:
SUMMARY.md
COMMANDS.md
LOG.md
SCIENTIFIC_CHANGELOG.md
COMPARABILITY_REPORT.md
status.json
ANNOTATED_README.md # original README + colored per-section agent-action annotations
PATCHES.md # only if patches were applied
Use the templates under assets/ and references/output-spec.md. Keep summaries
short, commands copyable, machine state stable, and scientific/comparability
changes explicit. ANNOTATED_README.md must preserve the source README byte-for-byte
outside inserted evidence blocks and pass its strip/check round trip. Use
--source-adjacent-readme only for an owned RIGORPILOT_README.md; never replace
an unrelated file. Distinguish verified facts from inference.
Reference Loading
- Workflow judgment:
references/agent-operating-principles.md. - Human-readable output:
references/language-policy.md. - Scientific/comparability judgment:
references/research-rigor-principles.mdand, when experiment details matter,references/deep-learning-experiment-principles.md. - Protocol-sensitive changes:
references/research-safety-principles.mdandreferences/patch-policy.md. - Personal rigor and lessons are advisory only; keep specialized detail in references/scripts rather than expanding this entrypoint.