Evals
Application evals — testing AI systems where the model is one component of a larger piece of software. Each entry is a repeatable technique for catching regressions, measuring accuracy, and finding the failures that create risk for the application and the people who depend on it. Not the foundational benchmarks that rank base models on leaderboards.
2 methods
Red team
Prompt-injection & jailbreak testing
Embed adversarial instructions in user input or retrieved/tool content to test whether the model follows attacker text over its system policy.
EstablishedPrompt injectionSafetyAgents
AdversarialSystem-prompt leakage testing
Try to get the model to reveal its own system prompt or hidden instructions, then check whether that prompt held anything — a credential, an access-control rule, unfiled business logic — that mattered only because nobody expected it to leak.
EstablishedPrompt injectionSafetyEU AI Act


