foundations-measurement-theory

v2026.09.24

Audit what scores and instruments measure. Use when auditing construct validity, measurement error, reliability, scales, calibration, invariance, or instrument drift.

GitHub
安装命令
npx skhub add vasilyu1983/foundations-measurement-theory
Markdown
SKILL.md

Measurement Theory Foundations

Establish whether an observation supports its intended interpretation and decision. A repeatable score can consistently measure the wrong thing.

When to use

Use when a task asks whether a metric, questionnaire, benchmark, sensor, rating, or composite score measures its intended target, or whether scores can be compared across groups, instruments, or time.

Do not use when the task is only sampling uncertainty, causal identification, choosing an action, or implementing an evaluation pipeline. Those belong respectively to statistical inference, causal inference, decision theory, and ai-evals.

Workflow

  1. Define the intended interpretation, population, decision, and cost of measurement error. Identify what is observed and what remains latent.
  2. Choose the physical-measurement or psychometric branch in measurement primitives. Do not translate psychometric reliability into metrological traceability, or treat benchmark scores as physical quantities without justification.
  3. Audit the instrument with the eight primitives below. Use evidence actually available; mark missing evidence rather than supplying thresholds or validity claims.
  4. For score comparisons, read comparability and drift. Preserve instrument versions, scoring changes, administration conditions, and population differences.
  5. Complete the measurement audit template, using the synthetic example only as a format example.

Quick Reference

Eight primitives:

PrimitiveRequired decision
Construct/measurandSpecify the target, domain, population, and intended use.
OperationalizationIdentify observation, instrument, scoring, and construct underrepresentation or irrelevant variation.
ValidityEvaluate evidence for this interpretation and use; never infer validity from reliability alone.
Reliability/precisionSpecify what is replicated: items, raters, occasions, tasks, or instruments; estimate the relevant uncertainty.
Scale/transformationsIdentify meaningful comparisons and operations; an arbitrary zero does not support ratio claims.
Error/uncertaintySeparate systematic bias, random error, missingness, and uncertainty in the measurement model.
Calibration/traceabilityIdentify reference and scope; calibration of a probability forecast is different from metrological calibration.
Invariance/driftDetermine whether the same interpretation survives group, time, setting, and instrument changes.

Completion criteria

Return a supported, conditional, or insufficient verdict for each intended interpretation, with evidence and restrictions. A verdict is an audit conclusion, not certification. Name the missing evidence and the smallest study or calibration needed to resolve it.

Do not use universal Cronbach alpha cutoffs. Internal consistency alone establishes neither unidimensionality, test-retest stability, agreement, nor validity. Keep item-level ordinal responses separate from assumptions used to analyze a composite.

Fact-Checking

Check source scope against the intended interpretation. Distinguish established definitions, empirical validation evidence, and synthetic illustrations. Sources were inspected through 2026-09-17; verify later standards or instrument changes before describing them as current.

Navigation

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

frameworks/shared-skills/skills/foundations-measurement-theory

默认分支

main

最新提交

8dc5de4

Tree SHA

700bf67