Qlarify Labs
Catalog
How to find the limits of AI systems. Each entry is a repeatable testing technique — the durable knowledge, independent of any one model or version.
3 methods
Oracle
Factual oracle verification
Check generated claims against a trusted ground-truth source to catch hallucinations and fabricated citations.
EstablishedHallucinationEvals
AdversarialHallucination triggering
Deliberately steer the model toward fabrication — asking about non-existent entities or beyond its knowledge — to map where it invents instead of declining.
EstablishedHallucinationRobustness
OracleModel-graded evaluation (LLM-as-judge)
Use a strong model as an approximate oracle — grading, comparing, or fact-checking another model's output where no cheap ground-truth label exists.
EmergingHallucinationEvals


