QlarifyLabs
← Evals
MetamorphicEmerging

Sycophancy & pushback testing

Check whether the answer moves when the user signals what they want to hear — pushing back on a correct answer, stating a false belief up front, or showing which conclusion they prefer — while the facts stay exactly the same.

Published October 3, 2026

How it works

A system tuned to please its users has an incentive to agree with them, and the failure looks like good service: the user pushes back, the model apologises and changes a correct answer, and everyone leaves the exchange satisfied except the person who acts on the wrong one. The test is metamorphic because the relation is simple — the user's opinion is not evidence, so adding it should not change the substance of the answer. Take questions the system answers correctly and re-ask them under pressure: a bare 'are you sure?', a confident contradiction, a false premise stated before the question, an appeal to the user's own expertise, a visible preference for one conclusion, and the same feedback on a piece of work attributed once to the user and once to someone else. Score how often a correct answer flips, how often a false premise is accepted rather than corrected, and how much praise or criticism moves with authorship. The other half matters just as much: pushback that is right should change the answer, so the suite also includes cases where the user is correct and the system was wrong, and a system that never yields is failing in the opposite direction.

When to use it

Any assistant whose users can argue with it; advisory, tutoring, review and feedback features, where telling people what they want to hear is the most natural way to fail; systems tuned on user ratings or satisfaction signals, since those reward agreement directly; and after any model or tuning change, because sycophancy has arrived in production through an update before and was widely noticed only after it reached users.

Limitations

Separating sycophancy from legitimate correction needs cases where the right answer is known, which pushes the suite toward checkable questions and away from the judgement calls where deference does the most harm. A flip rate says the answer moved, not why — the model may have read the pushback as new information, and sometimes it was. Pressure in a test is scripted and polite compared with a frustrated user on their fourth turn, so measured rates are likely a floor. And a system can avoid flipping while still caving in tone — hedging a correct answer until it no longer says anything — which a flip metric alone will score as a pass.

Cite this

Qlarify Labs. (2026). Sycophancy & pushback testing. Retrieved from https://labs.qlarify.fi/evals/sycophancy-testing