QlarifyLabs
← Evals
MetamorphicEmerging

Cross-lingual consistency testing

Ask the same question in each language the system serves and check the substance of the answer is the same — the facts, the refusal, the recommendation and the caveats, not just the topic.

Published October 3, 2026

How it works

A system offered in several languages is making one promise per language, and users in each assume they get the same product. They often don't: a model's factual knowledge and its refusal boundaries can both differ by language, and a system tested in English then shipped in Finnish or Swedish has, for those users, not been tested. The method is metamorphic: translating a request should not change the substance of the response. Build parallel sets — the same requests translated by people who know both the language and the domain, not by the model under test — and compare the answers across languages on what matters: the factual claims, whether the request was served or refused, which option was recommended, which warnings were included, and whether retrieval found the same source documents. Score agreement separately from accuracy, because the two can diverge — a system can be accurate in English and consistently wrong in another language, or inconsistent in a way that happens to average out. Mixed-language input belongs in the suite as well, since real users switch language mid-sentence, and so do cases where the source documents exist in only one of the languages.

When to use it

Any system offered in more than one language, especially where one is a smaller language with less training data behind it; customer-facing assistants in multilingual markets, where a different answer by language is a fairness problem as well as a quality one; safety and refusal behaviour, which is easy to test in one language and assume holds in the rest; and multilingual RAG, where the knowledge base and the question may not share a language.

Limitations

Parallel sets are only as parallel as the translation: a phrasing that is natural in one language can be odd in another, and part of what looks like inconsistency is the translation. Comparing open-ended answers across languages needs either a bilingual human or a judge model that is itself unevenly capable across those languages. Some legitimate differences belong in the answer — local law, local services, local conventions — so the suite has to mark where the answer should diverge. And consistency is not correctness: a system that is uniformly wrong in every language passes.

Cite this

Qlarify Labs. (2026). Cross-lingual consistency testing. Retrieved from https://labs.qlarify.fi/evals/cross-lingual-consistency-testing