Amira Ghazy ← Home
Game

Calibrate the judge

This judge arrives broken. Four unlabeled dials each introduce one fault. You're not told which is which — read the audit, find the dial behind each failing check, and clear all four. The fifth reading, agreement with humans, has no dial: it's the payoff that comes good on its own once the other four are in tolerance.

The dials

Move one, watch what changes, then bring it home.

The audit

Self-consistency≥ 0.80
Verbosity bias|r| ≤ 0.15
Group invariancegap ≤ 0.10
Position bias45–55%
Agreement w/ humansno dial · the payoff
No knob for this one. It climbs back on its own as the four faults above come into tolerance — agreement with the human criterion is what a reliable, unbiased judge gives you.

Same machinery as the audit demo, played backwards: there you add bias, here you remove it. The dials are verbosity, position, run-to-run noise, and a group tilt, scrambled and unlabeled. What evalaudit does in one function call, you're doing by hand.