Self-preference & cross-vendor bias testing
Measure whether a model that grades, routes to, or recommends other models favours its own family — by re-running the same comparison with authorship hidden, swapped, and deliberately mislabelled.
Published September 19, 2026
How it works
Systems increasingly put one model in charge of judging another: an LLM-as-judge scoring a competitor's answer, a router picking which model handles a subtask, an assistant recommending which SDK, format or tool to build on. Wherever that happens, the judge has a stake in the outcome, and the measured bias runs in its favour — models have non-trivial accuracy at recognizing their own output, and the stronger that self-recognition, the stronger the preference for it. The bias shows up in two directions that need separating: generosity toward its own family's output, and unwarranted harshness toward a rival's. A parallel version shows up outside scoring, in recommendation — defaulting to its own vendor's ecosystem when asked which library, model or platform to use, without the comparison that would justify it. The measurement is a metamorphic one, because the relation is simple: relabelling who wrote an answer should not change its score. Run a fixed set of response pairs through the judge blind, then labelled correctly, then with the labels swapped, then with the model's own output attributed to a rival — counterbalancing position at each step, so position bias doesn't get read as self-preference. The result is a win-rate delta per attribution condition, and a matching delta on ecosystem recommendations when the same question is asked with vendor names masked.
When to use it
Any evaluation harness where the judge model and a model under test come from the same family; mixed-vendor pipelines where one model reviews, routes, or hands off to another; before publishing comparative results produced with a model judge, since the judge's own family is the one result that needs a caveat; and for assistant features that recommend tools or platforms, where an unexamined default toward the vendor's own ecosystem is a product problem before it is an evaluation one.
Limitations
Measures a preference gap, not a motive — nothing here establishes intent, only that the verdict moved with the label. A gap can also be genuine: the judge's own family may really have produced the better answer, so separating favouritism from quality needs blinded human labels or an objectively scorable subset alongside the open-ended one. Blinding is never complete either, since style is a fingerprint the judge can read without a label, which makes the mislabelled condition the informative one rather than the blind baseline. And results are specific to a model-version pair and expire with the next release, so this is a standing check rather than a one-off audit.
Cite this
Qlarify Labs. (2026). Self-preference & cross-vendor bias testing. Retrieved from https://labs.qlarify.fi/evals/self-preference-bias-testing


