DifferentialBaseline-against-incumbent evaluation
Compare the AI feature with the process it replaces — the search box, the rules engine, the form, the person — on the same work, so 'it works' is measured against what users had before rather than against nothing.
EstablishedEvalsReliabilityProduction
OtherTrace-based production evaluation (AI observability)
Instrument the whole request — prompt, retrieval, model call, tool calls, tokens, latency, outcome — so real production behaviour becomes evidence you can query, sample, and score.
EmergingEvalsReliabilityProduction
BoundaryUnbounded-consumption / resource-abuse testing
Push volume, length, and concurrency past what the system enforces — unbounded prompts, recursive agent loops, or scripted request floods — to find where cost or capacity, not correctness, breaks first.
EstablishedReliabilityProduction
OtherUnit testing the deterministic scaffold
Test the deterministic code around the model — prompt builders, output parsers, schema validators, tool wrappers — in isolation, with exact assertions, the way you'd test any software.
EstablishedTool useReliability