evalaudit
A drop-in audit for LLM-judge and human-rating pipelines: reliability, agreement, bias, and measurement invariance in one call. The psychometric audit I kept writing by hand, packaged, tested, and documented.
Models, benchmarks, and LLM judges are measurement instruments. I treat them like one — auditing reliability and bias, packaging the methods into tooling, and writing the standards for reporting what an evaluation actually measures.
A drop-in audit for LLM-judge and human-rating pipelines: reliability, agreement, bias, and measurement invariance in one call. The psychometric audit I kept writing by hand, packaged, tested, and documented.
Treats an LLM-as-judge as a measurement instrument and audits it: test-retest reliability, verbosity and position bias, and whether it scores groups differently at equal quality. Includes a live, interactive demo.
An LLM-based scorer for open-ended responses — essays, interview answers, survey comments — validated the way an assessment would be: inter-rater reliability against human raters, score-distribution checks, and an adverse-impact analysis across demographic groups.
Drag sliders to give a judge the biases real evaluators have: verbosity, position, inconsistency, a group tilt. The measurement checks catch them live, on 300 simulated items in your browser.
You're the reviewer. Each round shows one judge's audit report and you call ship or flag. Eight rounds, then a score and a read on your eye.
Pick the better answer ten times, then get audited yourself: were you swayed by length, or by which answer you read first? The same biases I test models for, turned on you.
Targets show where an instrument's measurements landed. Read each pattern and call it — the bullseye game for the difference between precision and accuracy.
A judge arrives broken with four unlabeled dials. Diagnose which one drives each failing check and tune all five back into tolerance. The audit, played backwards.
Psychometrics spent a century learning how to measure things you can't see directly. AI evaluation keeps rediscovering the same lessons the hard way.
A short, fillable record of what an evaluation measures and what evidence backs it — a Model Card, but for the measurement validity of a benchmark or LLM judge. Several fields are answerable by running an audit, not by assertion.