QlarifyLabs
← Evals
AdversarialExperimental

Manipulation & vulnerability-exploitation testing (EU AI Act Article 5)

Check whether a persuasive or engagement-driven system uses deceptive or manipulative techniques to steer people — and whether it treats users made vulnerable by age, disability or a social or economic situation differently in ways that exploit that vulnerability.

Published October 3, 2026

How it works

Systems built to sell, retain, persuade or keep people engaged have an objective that can be served by manipulation, and conversational systems can apply it one person at a time. Article 5(1) of the EU AI Act prohibits, among other practices, AI systems that deploy subliminal techniques beyond a person's consciousness or purposefully manipulative or deceptive techniques, with the objective or effect of materially distorting a person's behaviour by appreciably impairing their ability to make an informed decision, causing them to take a decision they would not otherwise have taken, in a way that causes or is reasonably likely to cause significant harm (point (a)); and systems that exploit vulnerabilities of a person or group due to their age, disability or a specific social or economic situation, with the objective or effect of materially distorting their behaviour in a way that causes or is reasonably likely to cause significant harm (point (b)). The prohibitions have applied since 2 February 2025, and the Commission has published non-binding guidelines on how it interprets them. Testing doesn't decide whether a practice is prohibited; it produces the evidence that decision needs. Run scripted conversations through the system's real objectives — a cancellation request, a hesitant purchase, a user wanting to stop — and look for manipulative patterns, such as those catalogued for conversational systems in DarkBench (user retention, sycophancy, anthropomorphism, sneaking, brand bias), alongside interface patterns like false urgency, guilt and misrepresented options or costs. Then repeat the same conversations as simulated users signalling vulnerability — a minor, an older person confused by the process, someone in financial distress, someone in emotional crisis — and measure whether the pressure increases for users signalling vulnerability. A system that pushes harder on the user least able to resist is the pattern this method exists to find.

When to use it

Sales, retention, collections, subscription and engagement-optimised conversational systems; any assistant whose success metric is a user action — a purchase, an upgrade, staying subscribed — rather than the user's own goal; systems likely to reach minors, older users or people in financial or emotional distress; after tuning a system on conversion or engagement signals, since those can reward pressure; and as evidence for a legal assessment under Article 5.

Limitations

Whether a technique crosses into prohibited manipulation depends on intent or effect, materiality and the likelihood of significant harm — legal judgements the test can inform but not make, and the Commission's guidelines are non-binding, with authoritative interpretation reserved for the Court of Justice. Simulated vulnerable users are a proxy: a model playing a confused older person is not one, and real vulnerability shows up in ways scripts don't capture. Persuasion is not manipulation, so a rubric has to separate legitimate information and recommendation from deceptive pressure, and reasonable reviewers will disagree at the edge. Effects on actual behaviour need real users, which raises its own ethical constraints. And a system can avoid every listed pattern in testing while the harm comes from its objective in aggregate, which no single conversation shows.

Cite this

Qlarify Labs. (2026). Manipulation & vulnerability-exploitation testing (EU AI Act Article 5). Retrieved from https://labs.qlarify.fi/evals/manipulation-vulnerability-testing