OtherAI-transparency obligation testing (EU AI Act Article 50)
Check that the system meets the transparency duties the law places on its provider and deployer — people told they are talking to an AI, generated content marked as such — in every channel, at the latest at the first interaction, and every time rather than most of the time.
EmergingSafetyEvalsProduction
OtherAgent trajectory evaluation
Judge the whole path an agent took, not just its final message — did it reach the goal, in how many steps, with which tools, and did it loop, stall, or wander.
EmergingTool useEvalsAgents
DifferentialBaseline-against-incumbent evaluation
Compare the AI feature with the process it replaces — the search box, the rules engine, the form, the person — on the same work, so 'it works' is measured against what users had before rather than against nothing.
EstablishedEvalsReliabilityProduction
OtherModel & data supply-chain verification
Treat the base model, fine-tuning data, embeddings, plugins, and every third-party package pulled into the pipeline as a software supply chain — provenance-checked and version-pinned, not merely capability-tested.
EmergingSafetySupply chainEU AI Act
ExploratoryOff-scope & implicit-use testing
Find out what people actually use the system for, beyond what it was built for, and test how it handles those uses — serve them well, decline them clearly, or redirect — rather than improvising.
EmergingRefusalEvalsProduction
MetamorphicPerturbation testing
Apply small, meaning-preserving changes to an input — typos, spacing, paraphrase, reordering — and check that the output stays stable. When it doesn't, you've measured brittleness.
EstablishedReasoning failureRobustnessEU AI Act
MetamorphicSelf-preference & cross-vendor bias testing
Measure whether a model that grades, routes to, or recommends other models favours its own family — by re-running the same comparison with authorship hidden, swapped, and deliberately mislabelled.
EmergingBiasEvalsConsistency
Red teamSensitive information disclosure testing
Probe the model and its outputs for PII, credentials, or proprietary data it was never meant to reveal — memorized training data, leaked context, or a secret pasted earlier in the same conversation.
EstablishedSafetySensitive info disclosureEU AI Act
OtherTrace-based production evaluation (AI observability)
Instrument the whole request — prompt, retrieval, model call, tool calls, tokens, latency, outcome — so real production behaviour becomes evidence you can query, sample, and score.
EmergingEvalsReliabilityProduction
BoundaryUnbounded-consumption / resource-abuse testing
Push volume, length, and concurrency past what the system enforces — unbounded prompts, recursive agent loops, or scripted request floods — to find where cost or capacity, not correctness, breaks first.
EstablishedReliabilityProduction
OtherUnit testing the deterministic scaffold
Test the deterministic code around the model — prompt builders, output parsers, schema validators, tool wrappers — in isolation, with exact assertions, the way you'd test any software.
EstablishedTool useReliability