Amira Ghazy
Track 01

AI Evaluation & Measurement

Models, benchmarks, and LLM judges are measurement instruments. I treat them like one — auditing reliability and bias, packaging the methods into tooling, and writing the standards for reporting what an evaluation actually measures.

01

Selected Work

Python library · Psychometrics
Python
numpy · pandas
pip-installable

evalaudit

A drop-in audit for LLM-judge and human-rating pipelines: reliability, agreement, bias, and measurement invariance in one call. The psychometric audit I kept writing by hand, packaged, tested, and documented.

Demonstrates Turning measurement theory into reusable, tested tooling others can install and run.
Reliability · Bias Audit
Python
scikit-learn
psychometrics

Auditing the LLM judge

Treats an LLM-as-judge as a measurement instrument and audits it: test-retest reliability, verbosity and position bias, and whether it scores groups differently at equal quality. Includes a live, interactive demo.

Demonstrates Reliability and bias auditing of an automated evaluator, the way you'd validate any rating instrument.
Measurement · LLM Eval
Python
Anthropic API
scikit-learn
pandas

Does the model rate like a human?

An LLM-based scorer for open-ended responses — essays, interview answers, survey comments — validated the way an assessment would be: inter-rater reliability against human raters, score-distribution checks, and an adverse-impact analysis across demographic groups.

Demonstrates LLM engineering, psychometric validation (ICC, weighted kappa), and bias auditing in one pipeline.
02

Interactive

Live demo

Audit an LLM judge

Drag sliders to give a judge the biases real evaluators have: verbosity, position, inconsistency, a group tilt. The measurement checks catch them live, on 300 simulated items in your browser.

Game

Spot the broken judge

You're the reviewer. Each round shows one judge's audit report and you call ship or flag. Eight rounds, then a score and a read on your eye.

Game

Are you a biased judge?

Pick the better answer ten times, then get audited yourself: were you swayed by length, or by which answer you read first? The same biases I test models for, turned on you.

Game

Reliable, valid, both, or neither?

Targets show where an instrument's measurements landed. Read each pattern and call it — the bullseye game for the difference between precision and accuracy.

Game

Calibrate the judge

A judge arrives broken with four unlabeled dials. Diagnose which one drives each failing check and tune all five back into tolerance. The audit, played backwards.

03

Writing

Essay

AI Evaluation Is a Measurement Problem

Psychometrics spent a century learning how to measure things you can't see directly. AI evaluation keeps rediscovering the same lessons the hard way.

04

Standards

Reporting standard

The Validity Card for AI Evaluations

A short, fillable record of what an evaluation measures and what evidence backs it — a Model Card, but for the measurement validity of a benchmark or LLM judge. Several fields are answerable by running an audit, not by assertion.

People & Workforce Analytics → ← Back to home