OtherAI-transparency obligation testing (EU AI Act Article 50)
Check that the system meets the transparency duties the law places on its provider and deployer — people told they are talking to an AI, generated content marked as such — in every channel, at the latest at the first interaction, and every time rather than most of the time.
EmergingSafetyEvalsProduction
OtherAgent trajectory evaluation
Judge the whole path an agent took, not just its final message — did it reach the goal, in how many steps, with which tools, and did it loop, stall, or wander.
EmergingTool useEvalsAgents
DifferentialBaseline-against-incumbent evaluation
Compare the AI feature with the process it replaces — the search box, the rules engine, the form, the person — on the same work, so 'it works' is measured against what users had before rather than against nothing.
EstablishedEvalsReliabilityProduction
ExploratoryOff-scope & implicit-use testing
Find out what people actually use the system for, beyond what it was built for, and test how it handles those uses — serve them well, decline them clearly, or redirect — rather than improvising.
EmergingRefusalEvalsProduction
MetamorphicSelf-preference & cross-vendor bias testing
Measure whether a model that grades, routes to, or recommends other models favours its own family — by re-running the same comparison with authorship hidden, swapped, and deliberately mislabelled.
EmergingBiasEvalsConsistency
OtherTrace-based production evaluation (AI observability)
Instrument the whole request — prompt, retrieval, model call, tool calls, tokens, latency, outcome — so real production behaviour becomes evidence you can query, sample, and score.
EmergingEvalsReliabilityProduction