Qlarify Labs
Evals
How we evaluate AI systems to find their limits. Each entry is a repeatable testing technique — the durable knowledge, independent of any one model or version.
2 methods
Other
Bias auditing
Measure outcome disparities across a population of demographically varied inputs, reporting fairness as statistics rather than anecdotes.
EstablishedBiasEvals
MetamorphicCounterfactual bias probing
Hold a prompt fixed while swapping a protected attribute (name, gender, ethnicity) — the output should not change. When it does, you've measured bias.
EstablishedBiasEvals


