Declared-performance verification (EU AI Act Articles 13 & 15)
Re-measure the accuracy, robustness and per-group performance a provider declares for its system — under the conditions the declaration names — and check the system actually meets what its documentation says.
Published October 3, 2026
How it works
Documentation about an AI system makes claims: an accuracy figure, a metric, a list of conditions it was validated under, sometimes performance broken down by group. For high-risk systems the EU AI Act requires most of those claims — Article 15(3) requires the levels of accuracy and the relevant accuracy metrics to be declared in the instructions for use, Article 13(3)(b)(ii) requires the instructions to state the level of accuracy, including its metrics, robustness and cybersecurity against which the system has been tested and validated and which can be expected, and any known and foreseeable circumstances that may affect that level, and Article 13(3)(b)(v) requires, when appropriate, its performance regarding specific persons or groups of persons on which it is intended to be used. Where benchmark evaluation measures a system against a reference dataset, this method takes the system's own documentation as the reference. A declared number is a testable claim, and the method treats the documentation as the oracle. Rebuild the evaluation from the declaration: the same metric, computed the same way, on data that matches the stated conditions of use, and then on data from the actual deployment, where a gap is most likely to show. Check each declared group-level figure separately, since an aggregate that holds can hide a group that doesn't. Check the declaration's completeness as well as its truth — a figure with no stated metric, dataset or conditions can't be verified, and that is a finding in itself. Re-run it after every model, prompt or data change, because documentation can lag behind changes to the system.
When to use it
Providers preparing instructions for use for a high-risk system, before the figures are published; deployers adopting a third-party system, as an acceptance test of what the provider claims; after any change to the model, prompts, retrieval data or deployment population, to confirm the declaration still holds; and for any AI product whose marketing or documentation states performance figures, whether or not the Act applies.
Limitations
Only as good as the declaration is specific — a figure stated without its metric, test set or conditions can be neither confirmed nor refuted, which pushes the method toward auditing documentation quality as much as performance. Reproducing the provider's conditions may need data or access a deployer doesn't have. A system can meet every declared figure and still perform poorly on what the deployment actually needs, if the declaration measured the wrong thing. Per-group verification needs group labels, which raise their own data-protection questions. And whether a declaration meets the Act's requirements is a legal question; the test measures whether the system matches what it says, not whether what it says is sufficient. Under Regulation (EU) 2026/1744 (the Digital Omnibus on AI), the high-risk requirements apply from 2 December 2027 for Annex III systems and 2 August 2028 for systems covered by Annex I legislation; check the current consolidated text.
Cite this
Qlarify Labs. (2026). Declared-performance verification (EU AI Act Articles 13 & 15). Retrieved from https://labs.qlarify.fi/evals/declared-performance-verification


