Brand, persona & commitment testing
Check that the system speaks the way the organization wants to be seen — in voice, in role, and above all in what it promises — and never commits the organization to a policy, price or exception that doesn't exist.
Published October 3, 2026
How it works
A customer-facing assistant speaks for the organization, and what it says can be held against the organization. Testing that splits into three parts. Voice: does it keep the tone, register and vocabulary the organization uses, or drift into generic assistant phrasing, over-familiarity or sarcasm when a conversation turns difficult. Role: does it stay within the persona it was given, or can an ordinary user — not an attacker — talk it into commenting on competitors, politics or the company itself, writing content unrelated to its job, or disparaging its own product. Commitments: every statement about a policy, price, deadline, refund, warranty or exception is checked against the authoritative source, because a confident invented policy is the failure with legal weight — a small-claims tribunal has already found an airline liable for a bereavement-refund rule its chatbot stated and the airline's policy did not allow. The commitments part is an oracle test, against the policy documents rather than against what was retrieved, since a system can faithfully restate a stale document or fill a gap where nothing was retrieved at all. Run it with persistent, ordinary pressure: a customer who keeps asking for an exception, frames a request as urgent, or quotes what the bot 'said earlier'.
When to use it
Any assistant that talks to customers in the organization's name; sales, support and booking flows, where an invented exception has a direct cost; after any change to the system prompt, persona or knowledge base, since voice and policy accuracy both drift with them; and before launch, when the first screenshot of the bot saying something off-brand is still avoidable.
Limitations
Voice is a judgement call — a brand guide written for people rarely translates into a rubric precise enough for a judge model, so this part leans on human review and scales poorly. Commitment checking needs the policies written down, current, and unambiguous, which in many organizations they are not; the test then exposes the gap in the documentation as much as in the system. Ordinary-user pressure in a test is less inventive than real customers, and persona breaks found by deliberate attack belong to jailbreak testing, which this complements rather than replaces. And it can only check statements against policies someone thought to include — a promise about something no document covers passes unless a reviewer notices it shouldn't have been made at all.
Cite this
Qlarify Labs. (2026). Brand, persona & commitment testing. Retrieved from https://labs.qlarify.fi/evals/brand-persona-commitment-testing


