I measure what's hard to see — and prove the measurement holds.
A century of psychometrics worked out how to measure things that don't sit still — ability, bias, reliability, fairness. I bring that rigor to two places that badly need it: evaluating AI systems, and understanding the people inside organizations.
Approach
Most people who can build a model can't tell you whether its scores mean anything. Most people who understand measurement can't build the model. My work sits on both sides — and the same toolkit points at two domains.
Two tracks
AI Evaluation & Measurement
Treat models, benchmarks, and LLM judges as measurement instruments — and audit them like one: reliability, bias, and the validity evidence behind a score.
- — Auditing the LLM judge
- — evalaudit · a Python library
- — The Validity Card standard
- — Interactive demos, games & essay
People & Workforce Analytics
Measurement and causal rigor applied to the workforce — predicting attrition, then going past prediction to the people an intervention can actually keep.
- — Who should you actually keep?
- — Uplift modeling for retention
- — Attrition prediction & fairness
- — The Adverse-Impact Audit standard
Toolkit
Measurement & Statistics
- Reliability & validity theory
- Factor analysis · SEM · IRT
- Causal inference · uplift modeling
- Adverse-impact / fairness analysis
Machine Learning
- scikit-learn · XGBoost
- Model evaluation & interpretation
- SHAP · feature analysis
- Experiment design
LLM & AI
- Anthropic & OpenAI APIs
- LLM-as-evaluator pipelines
- Rubric design & human eval
- RAG · prompting
Engineering
- Python (pandas, NumPy)
- Git / GitHub
- Jupyter · reproducible analysis
- R (when the stats call for it)