Evals
Application evals — testing AI systems where the model is one component of a larger piece of software. Each entry is a repeatable technique for catching regressions, measuring accuracy, and finding the failures that create risk for the application and the people who depend on it. Not the foundational benchmarks that rank base models on leaderboards.
3 methods
Guardrail & moderation-layer testing
Treat the safety filters around the model as their own system under test — measuring bypass rate on unsafe inputs, false-positive cost on benign ones, and what happens when the guardrail itself fails.
Off-scope & implicit-use testing
Find out what people actually use the system for, beyond what it was built for, and test how it handles those uses — serve them well, decline them clearly, or redirect — rather than improvising.
Threshold testing
Walk inputs across a decision boundary — refusal, classification, confidence cutoff — to find exactly where the model's behaviour flips, and whether it flips in the right place.


