Evals
Application evals — testing AI systems where the model is one component of a larger piece of software. Each entry is a repeatable technique for catching regressions, measuring accuracy, and finding the failures that create risk for the application and the people who depend on it. Not the foundational benchmarks that rank base models on leaderboards.
3 methods
Boundary & edge-case testing
Push inputs to limits — very long contexts, token boundaries, empty/extreme values — where behavior tends to degrade sharply.
Instruction-following & constraint-adherence testing
Check that the system actually does what it was told — every constraint rather than most of them, still at turn forty, and with the right source winning when two instructions conflict.
Needle-in-a-haystack (long-context retrieval)
Plant a specific fact at varying depths in a long context and test whether the model can retrieve it from each position.


