Counterfactual bias probing
Hold a prompt fixed while swapping a protected attribute (name, gender, ethnicity) — the output should not change. When it does, you've measured bias.
Published August 22, 2026
How it works
A specialized metamorphic relation applied to fairness: hold a prompt fixed in every respect except a single demographic signal — a name that reads as more common to one gender or ethnicity, a pronoun, a stated age or nationality — and check that a fair system's output doesn't change. Systematically sweeping a matched set of names or pronouns across an otherwise-identical prompt (a résumé screen, a loan-application summary, a sentiment rating on identical text attributed to different authors) and comparing the distribution of outcomes quantifies differential treatment as a measured effect size, not an impression from a handful of anecdotes. Because only the demographic signal varies, a difference in outcome has no innocent explanation left to appeal to.
When to use it
Fairness audits before or after deploying any decision-support or evaluative use that touches people — hiring screens, lending and credit summaries, content moderation, sentiment or tone scoring, recommendation ranking.
Limitations
Requires careful construction of genuinely matched pairs — anything else that differs between the two prompts confounds the result. It only measures the specific attributes and value pairs actually tested, so a clean result on names is not evidence of fairness on, say, an inferred trait or an intersectional combination that was never probed.
Cite this
Qlarify Labs. (2026). Counterfactual bias probing. Retrieved from https://labs.qlarify.fi/evals/counterfactual-bias-probing


