Model-upgrade regression testing (history consistency)
Before moving a working system to a newer model, replay what it did on the old one and check the new version keeps every behaviour you rely on — better on average is not the same as no worse anywhere.
Published October 3, 2026
How it works
Where differential testing diffs two versions and drift monitoring watches a suite over time, this method focuses on the switch itself. A system built on one model version has an implicit specification: everything it currently does that users, prompts and downstream code have come to depend on. The old version's behaviour is the oracle: the expectation is that the product stays consistent with its own past. The method captures that behaviour as a replay set drawn from the system's own work — real or representative requests, the outputs that were accepted, and the properties of those outputs that something downstream relies on — then runs the same set, with the same prompts, tools and retrieval, through the candidate version before the switch. Two numbers matter, and the second is often skipped: the aggregate change, and the negative flips — individual cases the old version got right that the new one gets wrong. A model can score higher overall while losing cases that worked, and those lost cases are what users notice, because they had learned to rely on them. Improvement causes its own breakages too: a more capable model may be wordier, follow an instruction more literally, choose a different tool, format a field differently, or move a refusal boundary, and any of those can break a parser or a workflow while every quality score goes up. Because outputs vary from run to run, each case runs several times on both versions, and a wider spread on the new one counts as a regression even when its mean is higher. Prompts get tuned to the model they were written against, so the unit under test is the prompt-and-model pair: a workaround for the old version's quirks can be dead weight or actively harmful on the new one, and the comparison is only fair once the prompt has been reconsidered as well.
When to use it
Before adopting any new model version, including a minor one or a vendor's recommended replacement; when a provider announces a deprecation date, since the upgrade then happens whether or not anyone chose it; before switching providers or moving to a cheaper tier of the same family; after re-tuning prompts for a new model, to confirm the tuning didn't trade an old capability for a new one; and as the offline gate in front of a canary release, which can only catch what shows up quickly on live traffic.
Limitations
The replay set only protects behaviour someone thought to capture; what the system did incidentally and users quietly came to depend on is invisible until it breaks, so the set has to be fed continuously from production rather than written once. The old version is a reference, not ground truth — some of its accepted outputs were wrong, and a new version that diverges there is an improvement that the comparison will report as a regression, so every flip needs a human to decide which side was right. Exact-match comparison is useless for open-ended text and semantic comparison needs a judge model, which brings its own biases and is itself version-dependent. Separating the model's effect from the prompt's means running the old prompt and a re-tuned one, which doubles the cost. And the method answers whether the new version is safe to adopt, not whether it is worth adopting: a clean result on the cases you already had says nothing about what the new version could do that the old one couldn't.
Cite this
Qlarify Labs. (2026). Model-upgrade regression testing (history consistency). Retrieved from https://labs.qlarify.fi/evals/model-upgrade-regression-testing


