OtherAI-transparency obligation testing (EU AI Act Article 50)
Check that the system meets the transparency duties the law places on its provider and deployer — people told they are talking to an AI, generated content marked as such — in every channel, at the latest at the first interaction, and every time rather than most of the time.
EmergingSafetyEvalsProduction
DifferentialBaseline-against-incumbent evaluation
Compare the AI feature with the process it replaces — the search box, the rules engine, the form, the person — on the same work, so 'it works' is measured against what users had before rather than against nothing.
EstablishedEvalsReliabilityProduction
ExploratoryOff-scope & implicit-use testing
Find out what people actually use the system for, beyond what it was built for, and test how it handles those uses — serve them well, decline them clearly, or redirect — rather than improvising.
EmergingRefusalEvalsProduction
OtherTrace-based production evaluation (AI observability)
Instrument the whole request — prompt, retrieval, model call, tool calls, tokens, latency, outcome — so real production behaviour becomes evidence you can query, sample, and score.
EmergingEvalsReliabilityProduction
BoundaryUnbounded-consumption / resource-abuse testing
Push volume, length, and concurrency past what the system enforces — unbounded prompts, recursive agent loops, or scripted request floods — to find where cost or capacity, not correctness, breaks first.
EstablishedReliabilityProduction