Qlarify Labs
Evals
How we evaluate AI systems to find their limits. Each entry is a repeatable testing technique — the durable knowledge, independent of any one model or version.
Qlarify Labs
How we evaluate AI systems to find their limits. Each entry is a repeatable testing technique — the durable knowledge, independent of any one model or version.
22 methods
Serve two variants — prompts, models, or settings — to comparable slices of real traffic and let live outcomes decide which behaves better.
Deliberately craft inputs designed to elicit failure — confusion, unsafe output, or broken constraints — to map the model's weak boundaries.
Judge the whole path an agent took, not just its final message — did it reach the goal, in how many steps, with which tools, and did it loop, stall, or wander.
Score the model against a fixed, known-answer dataset so performance becomes a number you can track across versions and compare across models.
Measure outcome disparities across a population of demographically varied inputs, reporting fairness as statistics rather than anecdotes.
Push inputs to limits — very long contexts, token boundaries, empty/extreme values — where behavior tends to degrade sharply.
Test whether a model's stated reasoning actually determines its answer, or is a post-hoc rationalization.
Hold a prompt fixed while swapping a protected attribute (name, gender, ethnicity) — the output should not change. When it does, you've measured bias.
Run the same input across models or versions and treat divergence as a signal worth investigating.
Sample the model many times and test the distribution of its outputs — not any single answer — for drift, miscalibration, or instability.
Check generated claims against a trusted ground-truth source to catch hallucinations and fabricated citations.
Feed anomalous tokens, rare unicode, homoglyphs and malformed encodings to trigger out-of-distribution behavior.
Decompose an answer into its claims and check each one against the retrieved context — is the answer actually supported by what was retrieved, or did the model fill in the gaps?
Check that the model's outputs obey the rules of logic — valid inference, transitivity, symmetry, no self-contradiction — across related questions.
Test without an oracle by checking relations between the outputs of related inputs, instead of judging any single output in isolation.
Use a strong model as an approximate oracle — grading, comparing, or fact-checking another model's output where no cheap ground-truth label exists.
Specify invariants the output must always satisfy, then generate many inputs automatically and check the invariant on each.
Measure the retriever on its own terms — did the right documents come back, and how far up the ranking — before judging anything the model wrote with them.
Ask the same question multiple times (or multiple ways) and measure how often the answers agree.
A fast, shallow pass on every build — a handful of canonical prompts and health checks — whose only job is to fail loudly on gross breakage before anything deeper runs.
Deliberately probe whether a model abandons a correct answer or endorses a false one when the user pushes back, to catch approval-seeking behavior that plain accuracy evals miss.
Instrument the whole request — prompt, retrieval, model call, tool calls, tokens, latency, outcome — so real production behaviour becomes evidence you can query, sample, and score.