deterministic-metric-design

v2026.09.24

Inventing deterministic metrics — turning a fuzzy property like 'maintainability', 'risk', or 'how reducible this code is' into a deterministic, computable number an agent can trust and optimize. Covers the path from construct to adoption — operationalizing the construct, confronting computability limits (Kolmogorov, Rice) with sound proxies, picking the right measurement scale, proving properties (monotonicity, invariance, the Weyuker/Briand axioms), guaranteeing determinism, establishing construct validity (not just LOC in disguise), and hardening against Goodhart-style gaming when an agent optimizes the metric. Trigger when designing, reviewing, or validating a quantitative metric, score, measure, or index — and even when the user doesn't say 'metric' but wants to quantify, score, rank, or measure code/behavior, build a deterministic optimization target, or invent a measure for something previously unquantified (e.g., behavior-preserving codebase-size reduction).

GitHub
Install command
npx skhub add pproenca/deterministic-metric-design
Markdown
SKILL.md

dot-skills Deterministic Metric Design Best Practices

Design metrics that are deterministic, computable, provable, and valid — measures an agent can trust and optimize against without gaming them. The 44 rules across 8 categories take a metric from a fuzzy construct to an adoptable, machine-checkable number: define the construct, confront computability limits with sound proxies, ground it in measurement theory, prove its properties, pin its determinism, validate it empirically, harden it against optimization pressure, and package it for adoption.

A running example threads through every category — a deterministic measure of behavior-preserving codebase-size reduction (shrink code without changing how the app works). It is the ideal stress test because its ideal form is provably out of reach (Kolmogorov complexity is uncomputable; program equivalence is undecidable by Rice's theorem), so the whole craft is building a deterministic, tractable proxy with a proven guarantee.

This is the measurement-design layer that the *-algorithms skills apply (Big-O, NDCG, cyclomatic, MoJoFM) but never teach.

When to Apply

Use this skill when:

  • Designing a new metric, score, or index — or reviewing someone's proposed metric for rigor
  • Asked to "quantify", "measure", "score", or "rank" a property that has no agreed measure yet
  • Building a deterministic optimization target an agent will push on (e.g., reduce code size without changing behavior)
  • Auditing an existing metric that "feels off" — it suspiciously tracks LOC, jumps between runs, or gets gamed
  • Turning a research idea or formula into something computable, reproducible, and adoptable

Workflow: Define → Make Computable → Prove → Validate → Harden

The categories are ordered by cascade severity — an upstream mistake poisons everything below it. Work top-down, and jump straight to a category using this table:

If you are…Start inFirst rule
Starting from a fuzzy propertydef-def-name-the-latent-construct
Worried the ideal is uncomputable / undecidablecomp-comp-do-not-define-metric-as-uncomputable-ideal
Unsure whether you can average or take ratiosmeas-meas-declare-the-scale-type
Claiming the metric behaves a certain wayprop-prop-prove-monotonicity
Getting different numbers between runsdet-det-pin-iteration-and-tie-break-order
Unsure it measures the real thingvalid-valid-discriminant-not-just-loc
Letting an agent optimize the metricgame-game-hard-block-construct-violating-wins
Publishing the metric for othersagg-agg-ship-reference-impl-and-test-vectors

Each reference file is a {category}-{slug}.md containing: WHY it matters, an Incorrect example with the failure annotated, a Correct example with the minimal fix, and a reference. The incorrect/correct examples are metric definitions and procedures, not application code — the contrast is a badly-designed measure versus the fixed one.

Rule Categories by Priority

#CategoryPrefixImpactRules
1Construct Definition & Operationalizationdef-CRITICAL6
2Computability & Tractabilitycomp-CRITICAL7
3Measurement-Theoretic Foundationsmeas-HIGH5
4Proof of Metric Propertiesprop-HIGH6
5Determinism & Reproducibilitydet-HIGH5
6Construct Validity & Calibrationvalid-MEDIUM-HIGH6
7Optimization Safety & Anti-Gaminggame-MEDIUM5
8Aggregation, Reporting & Adoptionagg-LOW-MEDIUM4

See references/_sections.md for the full ordering rationale.

Quick Reference

1. Construct Definition & Operationalization (CRITICAL)

2. Computability & Tractability (CRITICAL)

3. Measurement-Theoretic Foundations (HIGH)

4. Proof of Metric Properties (HIGH)

5. Determinism & Reproducibility (HIGH)

6. Construct Validity & Calibration (MEDIUM-HIGH)

7. Optimization Safety & Anti-Gaming (MEDIUM)

8. Aggregation, Reporting & Adoption (LOW-MEDIUM)

How to Use

  1. Identify where you are with the Workflow table and open the matching first rule.
  2. Work the categories top-down — def- and comp- are CRITICAL because a fuzzy construct or an uncomputable ideal makes everything downstream noise or unusable.
  3. When proposing or critiquing a metric, quote the rule by file path so reviewers can check the reasoning.
  4. For a new metric, produce a one-page spec naming: construct, proxy, scale + unit + zero, proven properties, determinism guarantees, validity evidence, guardrails, and version — one line per category here.
  5. See references/_sections.md for ordering rationale and assets/templates/_template.md when adding rules.

Reference Files

FileDescription
references/_sections.mdCategory definitions, impact levels, and ordering rationale
assets/templates/_template.mdTemplate for adding new rules
metadata.jsonDiscipline, type, and source references

Related Skills

  • same-results-less-code, code-simplifier, complexity-optimizer, knip-deadcode — prescriptive code-reduction skills. This skill supplies the measurement layer they lack: a deterministic, behavior-preserving reduction metric to target and verify.
  • algorithmic-complexity-review, computer-science-algorithms — apply existing measures (Big-O). This skill teaches how to design new ones.
  • opensearch-function-scoring-algorithms — applied ranking metrics (NDCG, A/B tests). This skill is the foundational methodology beneath its eval- category.
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/.curated/deterministic-metric-design

Default branch

master

Latest commit

cf93c57

Tree SHA

afbb575