Amira Ghazy ← Home
Reporting standard · Draft

The Validity Card for AI Evaluations

A short, fillable record of what an evaluation actually measures — and what evidence backs the claim. Model Cards did this for models and Datasheets for datasets. Evals don't have one yet.

The gap

A benchmark score or an LLM judge's rating is a measurement, and measurements carry an implicit claim: that the number means what we say it means. Psychometrics spent a century formalizing how to defend that claim — the Standards for Educational and Psychological Testing (AERA, APA & NCME), Messick's unified validity, Kane's argument-based validation. Recent work has begun importing that apparatus into AI evaluation, framing a benchmark as a chain of inferences from response to real-world claim.

But that work lives in papers. There is no compact artifact a practitioner fills in — the way a Model Card (Mitchell et al.) or a Datasheet (Gebru et al.) travels with a model or dataset — that records the validity evidence for an eval and travels with it. This is a draft of that artifact: seven fields, each tied to an inference a score quietly depends on, several answerable by running an audit rather than by assertion.

The card

One card per evaluation (a benchmark, a rubric, an LLM-as-judge). A field left blank is not neutral — it marks an inference no one has defended.

01

Claim

intended interpretation & use
What construct does this eval claim to measure, and what decision rides on the score? Validity is a property of that interpretation and use, not of the test in the abstract.
02

Scoring

Kane: scoring inference
How is a response turned into a number — answer key, rubric, human raters, an LLM judge? What is the reliability of that scoring: test–retest across runs, inter-rater agreement?
evalaudit · self-consistency, agreement w/ humans
03

Generalization

Kane: generalization inference
Does performance on this item set generalize to the construct's domain? Coverage and sampling, item count, and the controls against contamination or leakage that would let a score generalize for the wrong reason.
04

Extrapolation

Kane: extrapolation inference · criterion validity
Does the score predict the real-world target it stands in for? Evidence linking eval performance to deployment outcomes — and an honest statement of the lab-to-production gap where that evidence is thin.
05

Invariance & fairness

measurement invariance · consequential validity
Does the instrument measure the same thing across groups and conditions, or does it tilt? Judge-side biases (verbosity, position), differential item functioning at equal true quality, and adverse impact on the decisions that follow.
evalaudit · verbosity, position, u