Baseline-against-incumbent evaluation
Compare the AI feature with the process it replaces — the search box, the rules engine, the form, the person — on the same work, so 'it works' is measured against what users had before rather than against nothing.
Published October 3, 2026
How it works
Most evals compare a model with another model or with a reference answer. The question that decides whether the feature should exist is different: is it better than what people did before it? Every AI feature replaces or supplements something — a keyword search, a decision tree, a template, a support agent, a manual review step — and that incumbent is the baseline that matters. The method runs the same set of real tasks through both and measures what the user experiences: task success, time to resolution, effort, error rate, how often a person had to step in, and cost per completed task. The incumbent is measured, not assumed — it is sometimes better than its reputation, and finding that a cheap deterministic baseline already solves most of the cases is a valid result. Results are broken down by case type and by user, because the average hides the pattern that usually matters: an assistant can help newcomers substantially while barely changing outcomes for experienced users, or win on easy cases and lose on the rare ones the old process handled well. In production this becomes a controlled rollout with a holdout group still on the old process.
When to use it
Before building an AI feature, with a prototype, to find out whether the incumbent leaves enough room to justify it; before retiring a non-AI process the feature is meant to replace; when the business case rests on a productivity or quality claim; and periodically afterwards, since the incumbent can improve too and the comparison that justified the feature can quietly stop holding.
Limitations
The incumbent is hard to run fairly once the new system exists — people stop maintaining the old search index or lose practice at the manual process, and the baseline degrades for reasons that have nothing to do with the AI. Outcomes that matter most, like customer trust or learning, show up slowly and are hard to attribute. A holdout group costs real users the better experience, if it is better. And the comparison is specific to the tasks sampled: a feature that wins on the common case can still be a worse replacement on the rare, high-stakes one, so case selection decides the verdict as much as the systems do.
Cite this
Qlarify Labs. (2026). Baseline-against-incumbent evaluation. Retrieved from https://labs.qlarify.fi/evals/baseline-against-incumbent-evaluation


