ai-evals

v2026.09.24

Designs trustworthy LLM, agent, responsible-AI, and multimodal evaluations. Use when measuring quality, fairness, privacy, grounding, safety, or judge reliability.

GitHub
安装命令
npx skhub add vasilyu1983/ai-evals
Markdown
SKILL.md

AI Evaluation and Fine-Tuning Methodology Skill

Core stance: an eval is an instrument. An untrusted instrument is worse than no instrument, because it produces confident wrong numbers that ship regressions. Fine-tuning is an optimization loop around that instrument. If the instrument is weak, training just makes the model better at gaming bad measurement. This skill is the cross-domain methodology layer that domain eval and model-lifecycle skills defer to: how to keep an LLM-as-judge honest, integrate eval frameworks, choose between prompting/context/tools/test-time compute/SFT/preference/RFT/PEFT/ distillation, derive thresholds instead of guessing them, and stop flaky runs or training leakage from masquerading as progress.

This is the methodology umbrella for evals. Domain skills own what to measure; this skill owns whether you can trust the measurement.

  • Building an eval system for a coding agent -> ai-coding-agents-observability-evals
  • Evaluating RAG / retrieval / search -> ai-rag
  • Running Hub model benchmarks (inspect-ai, lighteval) -> use the huggingface-skills: plugin (external)
  • General LLM lifecycle decisions -> ai-llm
  • This skill: judge bias, framework choice, calibration, reproducibility, optimization technique gates — the parts those four share and none owns in depth.

ASCII Flow

eval need
  |
  v
define verifiable goal  (what would FAIL if the requirement reverted?)
  |
  v
choose grader
  deterministic check  ->  LLM-as-judge  ->  human label  (cheapest that works)
  |
  v
control judge bias
  position / length / self-preference / verbosity
  |
  v
derive thresholds from a labeled calibration set  (not vibes)
  |
  v
choose optimization path  (prompt/RAG/tools -> SFT -> preference/RFT/PEFT)
  |
  v
choose inference-time lift  (self-consistency / rerank / verify / refine)
  |
  v
control flake  (pass@k, low temp, quarantine unstable cases)
  |
  v
trustworthy gate  ->  train / block / ship / rollback

Quick Reference

TaskRead or RunOutcome
Build a (question, ideal-answer) set and tune itreferences/dataset-construction.mdSourcing, ideal-answer authoring, run→compare→tune loop, robustness slice by perturbing memorized cases
Stop a judge from rating its own output highreferences/llm-judge-bias.mdSelf-preference, position, length, verbosity controls
Pick / wire an eval frameworkreferences/framework-integration.mdinspect-ai, lighteval, Ragas, DeepEval, promptfoo, Braintrust integration snippets + when to use each
Choose a pass threshold defensiblyreferences/threshold-derivation.mdDerive thresholds from a labeled set; inter-rater agreement; gate design
Stop flaky runs reading as regressionsreferences/flake-and-reproducibility.mdpass@k, seeds, temperature, quarantine, contamination/leakage, system-benchmark hazards (hardware lottery, thermal/background load, unreported CIs)
Decide if "A beats B" is real, size the setreferences/eval-statistics.mdBootstrap CIs, McNemar, power/MDE sizing, FDR, variance reduction
Get maximum from an LLMreferences/llm-optimization-technique-map.mdTechnique ladder across prompts, data, RAG/tools, test-time compute, SFT, preference/RFT, PEFT, distillation; measured case where a confidence scorer scored below no scorer
Decide whether and how to fine-tunereferences/fine-tuning-eval-loop.mdPrompt/RAG/tool baseline, SFT vs preference/RFT vs PEFT, split hygiene, promotion gates
Evaluate on live/production trafficreferences/online-production-eval.mdOffline-online correlation, A/B+guardrails, shadow/canary, drift, regression replay, HITL
Turn user behavior into eval and preference datareferences/conversational-feedback-signals.mdImplicit NL signals, action signals, edit→preference pairs, collection timing, feedback biases
Evaluate refusals, jailbreaks, harmreferences/safety-redteam-eval.mdOver/under-refusal, ASR per attack family, injection, harm rubrics, robustness
Evaluate responsible or multimodal AIreferences/responsible-multimodal-evaluation.mdIntersectional fairness, privacy/memorization/poisoning, oversight/appeals, provenance, grounding, diffusion, multimodal attacks, latency/cost
Go beyond one judgereferences/advanced-judging.mdJuries, fine-tuned judges, CoT/probability scoring, calibration (kappa/ECE), agentic reward

When to Use This Skill

Activate when the user asks for:

  • An LLM-as-judge / LLM grader and how to keep it honest
  • Why eval scores look inflated, noisy, or contradictory
  • Which eval framework to use, or how to integrate one
  • How to set or justify a pass/fail threshold or release gate
  • Pairwise / preference evaluation between two prompts, models, or harnesses
  • Reducing eval flakiness, contamination, or testset leakage
  • Calibrating a judge against human labels
  • Deciding whether to fine-tune, how to select SFT vs preference/RFT vs PEFT, or whether a fine-tuned model is genuinely better than a prompt/RAG/tool baseline
  • Getting maximum performance from an LLM using known techniques, including prompt/context/tool changes, test-time compute, reranking, distillation, or post-training
  • Testing fairness and intersectionality, differential privacy, explanations, poisoning, memorization, human oversight/appeals, watermarking/provenance, or environmental claims
  • Evaluating image-text retrieval, VQA, documents, video/audio, fusion, diffusion generation/control, multimodal safety, or multimodal latency and cost

Scope Boundaries (Use These Skills for Depth)

  • Domain metrics for retrieval (nDCG/MRR/recall, faithfulness) -> ai-rag
  • Agent golden tasks, tool-call grading, cost ops -> ai-coding-agents-observability-evals
  • Running benchmark harnesses on Hub models -> use the huggingface-skills: plugin (external)
  • Prompt CI/CD and structured output contracts -> ai-prompt-engineering
  • General model selection, serving, quantization, and deployment economics -> ai-llm and ai-llm-inference

Workflow

  1. Transform the vague ask into a verifiable goal. "Is it good?" is not gradeable. Ask: which case would fail first if the requirement reverted?
  2. Build the dataset before the grader. Source real questions, author ideal answers from the system's allowed context, and plan the run→compare→tune loop — see references/dataset-construction.md. No dataset, no eval.
  3. Pick the cheapest grader that works. Deterministic check > LLM judge > human. Reserve the LLM judge for what code cannot decide (Rule 5: use the model only for judgment calls).
  4. If using an LLM judge, control its bias before trusting any number — see references/llm-judge-bias.md. Untreated judge bias is the #1 source of confidently-wrong eval scores. When one judge isn't enough (high stakes, weak agreement, open-ended), escalate to juries / fine-tuned judges / calibrated scoring — see references/advanced-judging.md.
  5. Control flake and leakage with pass@k, low judge temperature, seed pinning, and held-out testsets — see references/flake-and-reproducibility.md.
  6. Derive thresholds from a labeled calibration set, not intuition or copied targets — see references/threshold-derivation.md. Size the gating set and judge "A beats B" with statistics (bootstrap CIs, McNemar, power/MDE, FDR) — see references/eval-statistics.md. For frozen paired results, run python3 scripts/analyze_paired_results.py results.csv --estimand unit_mean; declare whether units or clusters define the target population first. A population-level superiority claim needs design-compatible uncertainty; exact fixed-suite differences remain descriptive evidence.
  7. Only fine-tune after the baseline has earned it. Compare prompt/RAG/tool fixes first, then choose SFT for imitation/style/format/tool-call behavior, preference/RFT for rubric-scored reasoning or tradeoffs, and PEFT/LoRA/QLoRA when adapting an open model under compute or deployment constraints — see references/fine-tuning-eval-loop.md. Training loss is telemetry; held-out behavior is the verdict.
  8. Apply the full optimization ladder, not one pet method. For maximum LLM performance, evaluate cheap prompt/context/tool fixes, then inference-time methods (self-consistency, best-of-N, rerank/verify/refine), then data/SFT, preference/RFT/RLVR, PEFT, and distillation as the evidence warrants — see references/llm-optimization-technique-map.md. Each technique gets its own failure mode and gate.
  9. Gate loudly. A gate that passes while silently skipping cases is a failure dressed as success (Rule 12: fail loud). Report skipped/quarantined cases in the gate output.
  10. Extend past the offline gate where the system warrants it. Add safety/red-team evaluation (refusal precision/recall, jailbreak ASR, injection, harm rubrics) — see references/safety-redteam-eval.md — and, once in production, online evaluation (offline-online correlation, A/B with guardrails, drift, regression replay) — see references/online-production-eval.md. The offline gate is a filter; production is the verdict.

Core Principles

  • The judge is a model with failure modes. Treat its scores as one calibrated input, never as ground truth.
  • Different judge than the one under test. Self-preference bias is real and large; never gate on a model grading its own family/config.
  • Behavior, not plausibility. Rubrics that reward "looks good" reward length and confidence. Pin rubrics to verifiable behavior.
  • No threshold without a labeled set. A copied target (">95%") is a guess until validated on your own distribution.
  • No fine-tune without a baseline and a holdout. A tuned model that beats no prompt/RAG/tool baseline, or only wins on the training/dev set, has not earned release.
  • No "maximum performance" without a technique ladder. The best result often comes from composition: cleaner data + stronger retrieval/tool contracts + calibrated judge + small test-time search + selective post-training. Test the cheapest credible lift before moving weights.
  • Optimize behavior, not hidden knowledge. Fine-tune for stable formatting, domain style, tool-use patterns, rubric-following, or compact specialized behavior. Use retrieval/context for facts that change or must be cited.
  • Flake is a broken test, not a regression. A verdict that flips run-to-run means the eval is wrong, not the system.
  • Held-out or it's contaminated. If tuning ever saw the eval cases, the scores are inflated.
  • Goodhart's Law is the default outcome, not an edge case. Any metric that becomes a target (a threshold, a bonus, a promotion gate) will eventually be gamed — by the system under test, by whoever tunes against it, or by the judge itself. Every trap and anti-pattern in this file is a specific instance of this one law; treat a metric that stops correlating with the outcome you actually care about as expected decay, and re-anchor it against production outcomes or fresh human judgment on a schedule, not only when someone notices.
  • Match uncertainty to the claim. Population-level pass rates, win rates, and judge-human agreement need design-compatible uncertainty. Exact counts over a fixed deterministic suite describe that suite and need their denominator and limits; do not invent sampling intervals for them. See references/eval-statistics.md.

Release Decision Gate

Select foundations only when they change the eval decision

  • If a proxy score is being interpreted as quality, trust, fairness, or another construct, or compared across populations or grader versions, use measurement theory. Return the construct, observed proxy, interpretation evidence, comparability limits, and missing validation. Skip this audit for a routine execution of an already documented instrument with unchanged interpretation and population.
  • If uncertainty, dependent repeats/clusters, multiple comparisons, or repeated stopping can change a conclusion, use statistical inference. Return the estimand, independent unit, interval method and assumptions, and the multiplicity/stopping rule. Skip inference for reporting exact outcomes of a fixed deterministic fixture suite without a population claim; report its denominator and scope instead. Neither foundation substitutes for this skill's dataset, grader controls, or release decision.

Write the release rule before scoring: required tasks and slices, minimum acceptable effect, uncertainty method, regression budget, and handling for grader errors or missing results. A candidate passes only when every blocking slice is present and the predeclared decision rule clears. Learned, model, heuristic, and human graders must pass planted-good and planted-bad controls that detect calibration drift; deterministic exact-match, schema, executable-test, and invariant graders instead need versioned implementations with relevant positive and negative fixtures. Missing denominators, stale fingerprints, failed grader controls, or an unavailable judge produce inconclusive, never an implicit pass. Keep measured results separate from the release recommendation.

Known Traps

  • Grading an agent with the same model that produced the output (self-preference)
  • Comparing two candidates in fixed order and trusting the winner (position bias)
  • Copying a >95% threshold from a blog without validating it on your data
  • One judge call per request with no cheap deterministic pre-filter (cost blowup)
  • Treating a run-to-run verdict flip as a real regression instead of quarantining
  • Generating a synthetic testset from the same docs used to tune the system
  • Reporting "all passed" when some cases were skipped or errored (silent success)
  • Claiming "A beats B" from a point estimate with no confidence interval or test
  • Gating a small regression on a set far too small to detect it (no power check)
  • Tuning safety to block harm without a benign set, so the model over-refuses
  • Trusting a seed for reproducibility through a hosted API that isn't deterministic
  • Fine-tuning because the prompt is messy, the retrieval is broken, or the tool contract is ambiguous
  • Declaring the fine-tune better from training loss, validation loss, or one cherry-picked demo instead of a paired held-out eval with CIs
  • Letting the training set, grader calibration set, and release gate share cases
  • Training a judge or reward model on labels produced only by the same model family it will later grade

Common Anti-Patterns

  • Vibes-based eval: spot-checking a few outputs and calling it evaluation
  • Single-metric gates: one aggregate number hiding per-slice regressions
  • LLM-judge-only: no deterministic floor, so the gate inherits the judge's noise
  • Threshold-on-the-fly: setting the cutoff after seeing results to make it pass
  • Framework-as-strategy: adopting one vendor tool as the whole eval program
  • Fine-tune-as-strategy: reaching for SFT/RFT/LoRA before proving the failure is learned behavior rather than prompt, context, tools, or product spec
  • Technique soup: stacking CoT, self-consistency, rerankers, judges, and post-training without isolating which intervention caused the lift

Navigation

Resources:

Related skills:

Fact-Checking

  • Eval framework APIs (inspect-ai, lighteval, Ragas, DeepEval, promptfoo, Braintrust) change across releases. Verify current API and version against official docs before recommending a specific call or flag.
  • Fine-tuning platform support, model eligibility, dataset schemas, and RFT/grader APIs move quickly. Verify the current official docs before recommending a specific model, endpoint, hyperparameter, or CLI.
  • Optimization-method papers from arXiv are often preprint-only and benchmark-sensitive. Treat unreplicated methods as validate, not promote, until they beat a strong local baseline with cost/latency/safety gates.
  • Judge-bias findings (position, length, self-preference) are well-replicated through 2025-2026, but specific magnitudes are model- and prompt-dependent — re-measure on your own setup; do not quote a fixed number as universal.
  • Fine-tuned-judge models (Prometheus, JudgeLM, and successors) and jailbreak attack/defense results move fast — verify the current model, license, and benchmark-agreement claims before recommending a specific judge or asserting a model is robust to a given attack family.
  • Indirect prompt injection is the dominant agentic attack as of 2026; treat any "the agent is safe against injection" claim as requiring fresh adaptive testing.
  • Statistics methods (bootstrap, McNemar, FDR, power/MDE) are stable, but verify the exact API (scipy/statsmodels) before copying a call.
  • If web access is unavailable, mark framework-version claims as unverified.

Learnings Loop

Consult learnings.consolidated.md for relevant prior decisions or pitfalls; open learnings.md only when their history is needed. Skip both when unrelated to the task. After applying it, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py if you hit a pattern, mistake, or surprising fact. Do not modify SKILL.md itself.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

frameworks/shared-skills/skills/ai-evals

默认分支

main

最新提交

8dc5de4

Tree SHA

700bf67