sc-evaluate

v2026.09.25

LLM pipeline evaluation with oracle judge scoring. Runs prompts against gold standard datasets, evaluates output quality via LLM-as-judge, and generates scored reports with improvement recommendations.

GitHub
安装命令
npx skhub add tony363/sc-evaluate
Markdown
SKILL.md

LLM Evaluation Skill

Run LLM pipeline evaluation against gold standard datasets using oracle LLM-as-judge scoring. Measures output quality across weighted dimensions, identifies weak steps, and suggests prompt improvements.

Quick Start

# Full evaluation (all test cases, all steps)
/sc:evaluate

# Quick spot check
/sc:evaluate --cases=case_1,case_2 --steps=1,2,3

# Re-evaluate existing results without re-running pipeline
/sc:evaluate --skip-pipeline

# Generate outputs only (no evaluation)
/sc:evaluate --skip-eval

# Specify judge model
/sc:evaluate --judge-model=gpt-4o

# Dry run to preview plan
/sc:evaluate --dry-run

Behavioral Flow

  1. Discover - Find evaluation script, gold standards, and prompt files
  2. Configure - Parse scope (cases, steps, model overrides)
  3. Execute - Run pipeline on gold standard inputs
  4. Evaluate - Score outputs against gold standards via LLM-as-judge
  5. Analyze - Identify weak steps, dimension breakdowns, patterns
  6. Recommend - Suggest specific prompt improvements for low-scoring steps
  7. Report - Generate JSON + Markdown evaluation reports

Flags

FlagTypeDefaultDescription
--casesstringallComma-separated test case IDs to evaluate
--stepsstringallComma-separated step numbers to evaluate
--modelstringenv defaultOverride pipeline model
--judge-modelstringenv defaultOverride judge/oracle model
--skip-pipelineboolfalseSkip pipeline execution, evaluate existing results
--skip-evalboolfalseRun pipeline only, skip evaluation
--dry-runboolfalsePreview execution plan without API calls
--outputstringeval_runs/YYYYMMDD_HHMMSS/Output directory
--concurrencyint5Parallel judge calls
--thresholdint70Score threshold for "needs improvement"

Phase 1: Discover Project Structure

Locate evaluation components:

ComponentCommon LocationsPurpose
Evaluation scriptscripts/run_eval.py, eval/run.pyOrchestrates pipeline + scoring
Gold standardsgold_standards/, test_data/, fixtures/Expected outputs
Promptsprompts/, templates/Pipeline prompt templates
Rubricseval/rubrics.py, config/rubrics.yamlScoring dimensions and weights

If no standard structure found, ask the user to specify paths.

Phase 2: Configure Scope

Parse arguments to determine:

  • Which test cases to run (default: all discovered)
  • Which pipeline steps to evaluate (default: all)
  • Model overrides for pipeline and judge
  • Output directory (default: timestamped)

Create output directory:

OUTPUT_DIR="${output:-eval_runs/$(date +%Y%m%d_%H%M%S)}"
mkdir -p "$OUTPUT_DIR"

Phase 3: Execute Pipeline

Run the pipeline on gold standard inputs:

python <eval_script> \
  --output "$OUTPUT_DIR" \
  --verbose \
  [--cases CASES] \
  [--steps STEPS] \
  [--model MODEL] \
  [--skip-pipeline] \
  [--skip-eval]

API call estimation:

  • Pipeline: steps x cases API calls
  • Evaluation: scored_dimensions x cases judge calls

For quick validation, suggest running on 1-2 cases with 2-3 steps first.

Phase 4: Evaluate with LLM-as-Judge

For each step output, compare against gold standard using oracle LLM-as-judge:

Evaluation dimensions (customizable per project):

DimensionWhat It Measures
Content AgreementDo outputs cover the same key points?
Structure MatchIs the organization/format similar?
Detail AccuracyAre specific claims and data correct?
CompletenessAre all expected elements present?

Each dimension has a weight (0.0-1.0) summing to 1.0 per step.

Phase 5: Analyze Results

Read and analyze evaluation report:

  1. Overall similarity score across all cases and steps
  2. Per-step scores — highlight any below threshold (default: 70/100)
  3. Per-case scores — identify consistently weak test cases
  4. Dimension breakdowns for weak steps

Score interpretation:

Score RangeAssessmentAction
85-100ExcellentNo changes needed
70-84GoodMinor tuning possible
60-69Needs improvementPrompt revision recommended
Below 60PoorPrompt likely needs rewrite

Phase 6: Recommend Improvements

For each step scoring below threshold:

  1. Read the current prompt template
  2. Read the gold standard output (expected)
  3. Read the pipeline output (actual)
  4. Compare and identify gaps:
    • Missing instructions that gold standard captures
    • Overly broad instructions causing divergent output
    • Format/structure differences
    • Specificity gaps

Present actionable suggestions:

### Step N: <step_name> (Score: XX/100)

**Weakest Dimension**: <dimension> (XX/100)

**Gap Analysis**:
- Gold standard includes <X> but prompt doesn't instruct it
- Output format diverges: gold uses <format>, output uses <other>

**Suggested Prompt Changes**:
1. Add instruction: "<specific instruction>"
2. Clarify format: "<format guidance>"
3. Add example: "<example output snippet>"

Output Structure

eval_runs/YYYYMMDD_HHMMSS/
  results/                    # Pipeline outputs
    case_1/
      step_01_<name>.md
      step_02_<name>.md
      ...
    case_2/
      ...
  evaluation/                 # Judge scores
    evaluation_report.json
    evaluation_report.md
    per_step_scores.csv
    per_case_scores.csv

MCP Integration

PAL MCP (Optional)

ToolWhenPurpose
mcp__pal__thinkdeepLow-scoring stepsDeep analysis of why outputs diverge
mcp__pal__consensusPrompt revisionMulti-model validation of proposed changes
mcp__pal__codereviewEval scriptReview evaluation pipeline code

Rube MCP (Optional)

ToolWhenPurpose
mcp__rube__RUBE_REMOTE_WORKBENCHLarge eval runsProcess results in Python sandbox
mcp__rube__RUBE_MULTI_EXECUTE_TOOLNotificationsReport results to Slack/email

Error Handling

ScenarioAction
No eval script foundAsk user for script path
No gold standards foundAsk user for gold standard directory
API rate limitReduce concurrency, add delays
Pipeline step failsLog error, continue with remaining steps
Judge returns invalid scoreRetry once, then flag for manual review
Output directory existsAppend timestamp suffix

Guardrails

  • Always pass --verbose for progress visibility
  • Warn about API call counts before full runs
  • Suggest quick validation on subset before full evaluation
  • Preserve all intermediate outputs for debugging
  • Never modify gold standard files

Tool Coordination

  • Bash - Run evaluation scripts
  • Read - Inspect prompts, gold standards, outputs, reports
  • Write - Generate reports
  • Grep - Search for patterns in outputs
  • PAL MCP - Deep analysis of score gaps
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.25

发布时间

2026年9月25日

分类

未分类

许可证

MIT

源路径

.claude/skills/sc-evaluate

默认分支

main

最新提交

6634f8e

Tree SHA

993acfd