QlarifyLabs
← Evals
OtherExperimental

Feedback-loop testing (EU AI Act Article 15(4))

For systems whose outputs feed back into their future inputs — retraining on their own decisions, retrieving their own generated content, ranking by engagement they produced — simulate the loop over many cycles and measure whether error or bias amplifies.

Published October 3, 2026

How it works

Some AI systems change their own future: a ranking system's choices decide which items get the clicks it later learns from; a risk model's flags decide which cases get investigated and so which outcomes are ever observed; an assistant's generated answers end up in the knowledge base it retrieves from. In each case a small bias in the output can become a larger bias in the data, and the system drifts in a direction no single evaluation would show. Article 15(4) of the EU AI Act names this for high-risk systems that continue to learn after being placed on the market or put into service: they must be developed so as to eliminate or reduce as far as possible the risk of possibly biased outputs influencing input for future operations (feedback loops), and to ensure that any such loops are duly addressed with appropriate mitigation measures. Testing it means running the loop deliberately rather than waiting for it. Map where outputs re-enter as inputs — training labels, retrieval corpora, user-behaviour signals, which cases get human review. Then simulate many cycles against a known ground truth: let the system's decisions determine what data it sees next, retrain or re-index as production would, and track accuracy, per-group outcome rates and the diversity of what it surfaces over cycles. A loop is present when the distance from ground truth grows with cycles rather than staying flat, and a loop that only shows up in one group or one region of the input space is still a loop. Repeat the simulation with whatever mitigations the system uses in place, to check whether the amplification still appears.

When to use it

High-risk systems that continue to learn after deployment, where Article 15(4) applies; recommender, ranking and allocation systems whose outputs decide what data is collected next; RAG systems that write generated content back into their own corpus; any system retrained on labels that depend on its own past decisions; and before turning on continuous learning for a system that was previously retrained only offline.

Limitations

Simulation needs a model of how the world responds to the system's decisions — which users click, which cases get investigated — and the result is only as good as that model; a real loop can run through a behaviour the simulation didn't include. Ground truth is exactly what a feedback loop corrupts, so the test needs data the system never influenced, which many deployments no longer have. Loops may be slow, appearing over many production cycles that a simulation compresses imperfectly. Distinguishing a harmful loop from legitimate adaptation to a changing world needs a judgement about what the system should be tracking. And the method shows whether amplification happens, not whether the mitigations satisfy the Act, which is a legal question. Under Regulation (EU) 2026/1744 (the Digital Omnibus on AI), the high-risk requirements apply from 2 December 2027 for Annex III systems and 2 August 2028 for systems covered by Annex I legislation; check the current consolidated text.

Cite this

Qlarify Labs. (2026). Feedback-loop testing (EU AI Act Article 15(4)). Retrieved from https://labs.qlarify.fi/evals/feedback-loop-testing