Bias auditing
Measure outcome disparities across a population of demographically varied inputs, reporting fairness as statistics rather than anecdotes.
Published August 22, 2026
How it works
Where counterfactual probing swaps one attribute in a matched pair, a bias audit zooms out to the population: run many demographically labelled inputs through the system — resumes with names varied by inferred ethnicity, support tickets with gender-coded phrasing swapped — and measure outcome rates, error rates, and sentiment across groups, the way a fair-lending or hiring audit would. The output is a set of disparity metrics with confidence intervals per group, not a single anecdote: an approval-rate gap of a few points on a hundred paired examples is noise, the same gap across ten thousand is a finding. That statistical form is both what holds up under scrutiny and the form regulators increasingly expect.
When to use it
Any evaluative or decision-support use touching people — screening, moderation, recommendation, sentiment; especially where there is legal exposure; and worth running before any protected-attribute-adjacent feature reaches production, not only after a complaint surfaces one.
Limitations
Only audits the attributes and groups you chose to measure; intersectional and unmeasured biases slip through, and a clean audit on tested dimensions is not a clean bill of health. Disparity metrics also need a large enough sample per group for the confidence intervals to mean anything, which is easy to underestimate for smaller subgroups.
Cite this
Qlarify Labs. (2026). Bias auditing. Retrieved from https://labs.qlarify.fi/evals/bias-auditing


