Reporting standard · Draft
The Adverse-Impact Audit
A bias-audit standard for automated employment decision tools — hiring, ranking, promotion, attrition flags. The four-fifths ratio is where the law starts. It is not where the bias is.
The gap
Bias audits of hiring algorithms are now mandated, not optional — New York City's Local Law 144, Colorado's SB 24-205, the EU AI Act's high-risk classification of employment tools. Most audits answer the question the law asks literally: compute a selection-rate ratio, check it against the four-fifths rule, publish it.
That ratio is a screen, not a diagnosis. It tells you a group is selected less often; it cannot tell you whether the tool is wrong about them. The deeper question — does the tool predict a group's actual job performance fairly — is the one industrial-organizational psychology has answered since the 1970s, through the Uniform Guidelines on Employee Selection Procedures and Cleary's regression model of test bias. That layer is largely absent from the tooling being sold today. This is a draft standard that records both: the legal floor, and the predictive-bias evidence underneath it.
The audit
One record per tool-and-decision. The fields run from the legal floor upward; the serious findings usually live below the ratio, not in it.
01
Tool & decision
scope
What is the tool, what employment decision does it drive — screen, rank, score, recommend, flag — and over what applicant or employee population? An audit is scoped to a decision, not to a model in the abstract.
02
Adverse impact
four-fifths rule · statistical significance
Selection or scoring rate for each protected group against the highest-selected group: the impact ratio read against the four-fifths (80%) threshold, by sex, race/ethnicity, and intersectional category — plus a significance test, since the ratio alone is unstable at small N.
03
Differential validity
predictor–criterion correlation by group
Does the score predict the job-relevant outcome equally well across groups? Equal validity is necessary but not sufficient — a tool can correlate with performance equally in two groups and still be biased in how it predicts.
04
Differential prediction
Cleary (1968) regression model of test bias
Regress the real outcome on the tool's score, group, and their interaction. A common prediction line that systematically over- or under-predicts a group's actual performance — an intercept or slope difference — is predictive bias in the legal and scientific sense. The layer most audits skip.
evalaudit · uniform & non-uniform DIF (same regression engine)
05
Measurement invariance
metric & scalar invariance
Independent of outcomes, does the instrument measure the same construct across groups — equal loadings (metric) and equal intercepts (scalar)? Scalar non-invariance is the measurement-side mirror of uniform predictive bias.
06
Job-relatedness & alternatives
Uniform Guidelines · business necessity
Where adverse impact exists, is the tool validated as job-related and consistent with business necessity, and was a less-discriminatory alternative of comparable validity sought and documented? The duty to look for that alternative is part of the standard, not a courtesy.
07
Consequences & cadence
governance
Who bears a false screen-out, what is disclosed to candidates, and how often is the tool re-audited — at minimum annually and on every model or data change? Maps to Local Law 144's published impact ratios and Colorado SB 24-205's risk-assessment and documentation duties.
Why it goes past the ratio
The four-fifths ratio is the floor the law sets. Differential prediction is where the science lives — and the two can disagree. A tool can clear the ratio and still under-predict a group's real performance; it can fail the ratio for reasons that turn out to be job-related. An audit that reports only the ratio answers the wrong question precisely.
Grounded in the Uniform Guidelines on Employee Selection Procedures (EEOC, 1978), Cleary's regression model of test bias (1968), and the Standards for Educational and Psychological Testing — applied to the 2023–2026 wave of AEDT law. The differential-prediction and invariance fields run on the same regression engine as evalaudit and the model in attrition-fairness. Sibling to the Validity Card for AI evaluations: the same measurement logic, pointed at employment decisions instead of model scores.