Amira Ghazy
I/O Psychology · Applied AI · People Analytics

I measure what's hard to see — and prove the measurement holds.

A century of psychometrics worked out how to measure things that don't sit still — ability, bias, reliability, fairness. I bring that rigor to two places that badly need it: evaluating AI systems, and understanding the people inside organizations.

Calibration reading: the intersection
AI systems The workforce
01

Approach

Most people who can build a model can't tell you whether its scores mean anything. Most people who understand measurement can't build the model. My work sits on both sides — and the same toolkit points at two domains.

03

Toolkit

Measurement & Statistics

  • Reliability & validity theory
  • Factor analysis · SEM · IRT
  • Causal inference · uplift modeling
  • Adverse-impact / fairness analysis

Machine Learning

  • scikit-learn · XGBoost
  • Model evaluation & interpretation
  • SHAP · feature analysis
  • Experiment design

LLM & AI

  • Anthropic & OpenAI APIs
  • LLM-as-evaluator pipelines
  • Rubric design & human eval
  • RAG · prompting

Engineering

  • Python (pandas, NumPy)
  • Git / GitHub
  • Jupyter · reproducible analysis
  • R (when the stats call for it)

Open to work in AI evaluation, people analytics, and psychometric data science.