QlarifyLabs
← Evals
ExploratoryEmerging

Off-scope & implicit-use testing

Find out what people actually use the system for, beyond what it was built for, and test how it handles those uses — serve them well, decline them clearly, or redirect — rather than improvising.

Published October 3, 2026

How it works

A system built for one job is used for many. For example, a shopping assistant gets asked for medical advice, a code helper is used to draft a complaint letter, an internal knowledge bot becomes the place people ask HR questions it has no documents for. None of these are attacks, and none show up in a test suite written from the product spec, so the system meets them for the first time in production and improvises. The method has two halves. Discovery: mine real usage — production traces, support tickets, logs from a pilot — for clusters of requests the system wasn't designed for, and rank them by how often they occur and how much harm a bad answer would do. Testing: for each cluster, decide what the right behaviour is — serve it, decline it, or hand off — and then test that behaviour like any other requirement, including how a decline is worded and whether it points somewhere useful. A clear 'I can't help with that, here's who can' is a designed outcome; a confident answer in a domain the system has no grounding in is the failure this method exists to find.

When to use it

After any launch, once real traffic exists to mine; general-purpose interfaces — a chat box with no stated limits invites everything; systems where an out-of-scope answer could cause harm, such as health, legal, financial or safety-adjacent questions arriving at a system built for something else; and before deciding whether to extend a product's scope, since the implicit uses are the evidence of what people wanted it to do.

Limitations

Discovery depends on having usage data and the right to analyse it, which privacy constraints can limit, and it only finds uses that already happened — not the ones a launch in a new market will bring. The right behaviour for each cluster is a product decision, not a test result, and teams often avoid making it, leaving the system to improvise by default. Over-correction is its own failure: a system tuned to decline off-scope requests starts refusing legitimate in-scope ones that merely look unusual, so in-scope pass rates have to be watched alongside. And the long tail never ends; this is a standing review of what changed in usage, not a suite that is ever finished.

Cite this

Qlarify Labs. (2026). Off-scope & implicit-use testing. Retrieved from https://labs.qlarify.fi/evals/off-scope-implicit-use-testing