How to read the catalog
Our approach
We are not collecting AI bugs as trophies. We are studying how to put AI to the test — and cataloging the techniques that do it, independent of any one model or moment.
AI moves fast. A model that fails a technique today may be patched next week, and a sharp piece of writing can read as dated within a season. If the goal were a live scoreboard of current bugs, that churn would be fatal. It is not our goal.
The durable thing — the part worth citing a year from now — is the method: the repeatable technique that surfaces a class of failure no matter which model you point it at. Read this site that way and the pieces fall into place.
The method is the point.
A specific bug is an artifact of one model version at one moment — patched, it loses its shine. The technique that surfaced it is durable, transferable knowledge: it still works on the next model, the next version, the next vendor. That's what the catalog holds — repeatable methods for finding where AI systems fail, not a scoreboard of individual failures.
Maturity is stated, not implied.
Every method says whether it's an established discipline — regression suites, differential and metamorphic testing, property-based checks — or a newer, less-proven technique. We don't dress up the experimental as settled, and we don't bury the boring-but-reliable behind the flashy.
Resources are dated and rated, not just linked.
AI moves fast enough that a sharp piece of writing can read as dated within a season. Every external resource we curate carries a credibility rating and, where we can find it, a publish date — so you can judge how much to lean on it today, not just whether it was good when it was written.
A real testing practice blends novel and established methods.
The flashy adversarial prompt gets the headlines, but a serious testing harness leans just as hard on the boring, reliable disciplines. Established methods earn their place precisely because models regress — we give them equal billing, not an afterthought tacked on behind the jailbreaks.
Why we are vocal about this
Plenty of AI content is written as if each result were a permanent truth. In a field that re-bases itself every few months, that ages badly and quietly misleads. We would rather state the obvious out loud: this is a study of testing methods, maintained with the assumption that models will keep changing under us. That stance is the whole point — so we are putting it on its own page.


