Est.

EU AI Act Requirements for Agentic AI Systems

Enterprises must map agentic AI to existing Act rules without a special regulatory category.

Correspondent · · 11 min read
Cover illustration for “EU AI Act Requirements for Agentic AI Systems”
MCP Policy and Compliance · September 15, 2026 · 11 min read · 2,530 words

The EU AI Act never uses the phrase "AI agent." Not once, in the whole regulation. Yet Regulation (EU) 2024/1689 already governs agentic systems in full, because the definitions it does use, an AI system under Article 3(1) and a general-purpose AI model under Article 3(63), are broad enough to catch them without needing a special category. That's the Commission's own reading, laid out in the AI Act Service Desk FAQ. So the absence of a label doesn't mean an absence of obligation. It means enterprises have to do the mapping work themselves.

For this piece, an "agent" means what Nannini and colleagues described in their April 2026 paper: an AI system that plans on its own, calls external tools, and runs multi-step chains of action with a person only loosely in the loop. That has moved beyond a lab concept. Agents already screen job candidates, support clinical decisions, manage pieces of critical infrastructure, and handle customer service at scale. The Future Society's June 2025 report, "Ahead of the Curve: Governing AI Agents under the EU AI Act," was the first real attempt to work through what that means under the Act, and it confirmed the coverage while flagging real gaps that still need Commission guidance and updated technical standards.

The upshot for compliance teams: you can't sort agents by category and call it done. Every use case has to get its own read, because the words "AI agent" carry no risk tier on their own. The Act works across three layers at once, prohibitions, GPAI model duties, and high-risk system obligations, and a single agentic deployment can trip all three simultaneously.

How the Act's tiered structure maps onto agentic deployments

The Act sorts everything into four tiers: unacceptable risk (banned outright), high risk (full obligations), limited risk (just transparency duties), and minimal risk (voluntary codes, nothing mandatory). Most agentic deployments in the enterprise world land in the high-risk or limited-risk bands. Minimal risk is reserved for genuinely low-stakes stuff, summarizing a document, drafting a first pass at an email. Nothing that touches a hiring decision or a loan application.

Per the Salt Security analysis, the Annex III categories that keep catching agents are hiring and HR, credit decisions, critical infrastructure, law enforcement, and essential public services. And here's the nuance The Future Society flagged in June 2025 that trips people up: an agent built to work across multiple domains gets assumed high-risk by default, unless the provider can show it took real precautions to prevent that. Building something "general-purpose" is not a workaround for Annex III. It's often the opposite, a faster route into it.

There's a second layer running underneath all this. The Act regulates at the AI system level, the application most enterprises actually deploy, and separately at the GPAI model level, the foundation model sitting underneath that application. A model only becomes a "system" once you bolt something onto it, a user interface, say, per Recital 97. That distinction matters a lot for anyone fine-tuning or wrapping a foundation model for their own use. Recitals 99 and 100 spell out the same idea from the other direction: large generative models are the textbook example of a general-purpose AI model, and once one gets embedded into a system, that system counts as general-purpose too, capable of serving all sorts of downstream purposes its original trainers never scoped out.

Practically, this means a company running a pipeline of agents that includes one screening job applicants cannot treat compliance as confined to that single agent alone, the broader deployment context shapes the obligation.

What the enforcement timeline actually looks like after the Digital Omnibus

Diagram: The EU AI Act Enforcement Timeline for Agentic Systems. Visualizes: Show a linear timeline of the key enforcement dates under the EU AI Act as they apply to agentic deployments.

Article 5's prohibitions have been in force since February 2, 2025. No phase-in period, no exception carved out for agentic architectures because they're new or complicated. Penalties under Article 99(3) became enforceable August 2, 2025: up to €35 million or 7% of worldwide annual turnover, whichever number is bigger. GPAI model obligations under Articles 53 and 55 also started August 2, 2025. Article 50's transparency rules, chatbot disclosure, deepfake labeling, kick in August 2, 2026, and the Digital Omnibus left that date untouched.

The Digital Omnibus itself is worth understanding, because it's easy to misread as a broad delay. The Commission published its proposal on November 19, 2025, prompted mainly by member states falling behind on naming national competent authorities and finalizing harmonized technical standards. Regulation (EU) 2026/1744 pushed the Annex III high-risk obligations from August 2, 2026 out to December 2, 2027, a 16-month extension. Annex I embedded high-risk systems got pushed even further, to August 2, 2028.

The Omnibus also added a new absolute prohibition: AI systems that generate or manipulate non-consensual intimate imagery, or CSAM, are banned outright, with that ban taking effect December 2, 2026 after a transition window. A separate transitional watermarking rule under Article 50(2) runs on the same December 2026 date for systems already on the market before August 2026.

The Commission is aiming for at least a 25% cut in compliance and documentation costs overall by 2029, and 35% for small and mid-sized companies. That's the stated goal, not a guarantee. More than 60 civil society organizations pushed back on the negotiated text, warning it weakens fundamental-rights protections, and pointed out that no full impact assessment ran before publication.

Here's what the extension does not touch: Article 5 prohibitions, GPAI obligations, and transparency requirements are all still running on their original clocks. Any enterprise treating December 2027 as the start line for compliance work is already behind on three separate fronts. And incident reporting under Article 73 is active right now: 24 hours for life- or safety-threatening incidents, 72 hours for other serious incidents.

The prohibitions already binding on every agent deployment

Article 5 lists eight prohibited practices. Seven are flat bans, no exceptions. The eighth, real-time remote biometric identification in public spaces, has a narrow carve-out for law enforcement, and even that requires prior judicial or administrative sign-off.

The eight: manipulative or subliminal techniques, exploiting vulnerabilities, social scoring, individual criminal risk assessment based purely on profiling, untargeted scraping of faces for recognition databases, emotion recognition at work or in schools, biometric categorization tied to sensitive traits, and real-time remote biometric ID for law enforcement. These bind providers and deployers both, each within their own slice of responsibility, per the Commission's Service Desk guidance.

For agentic systems specifically, two of these deserve the most attention: harmful manipulation under Article 5(1)(a), and exploitation of vulnerabilities under 5(1)(b). Neither gets satisfied with a policy document sitting in a compliance folder somewhere. They need to be designed for, at the architecture level. An agent that adapts its persuasion tactics based on a user's inferred mood, or personalizes pressure based on someone's financial stress signals, sits squarely in this territory. Agents deployed in HR, benefits administration, or consumer financial services need a specific Article 5 review, not a general one.

The new non-consensual intimate imagery prohibition added by the Digital Omnibus matters directly for multimodal generative agents, anything that can produce or alter images or video as part of its output.

Because Article 5 is live now, not in 2027, waiting for Annex III guidance to sort this out isn't an option. Any agent that talks to users, makes a recommendation, or nudges a decision needs an Article 5 clearance review today, not on some future compliance roadmap.

GPAI model obligations and what they mean for the agent layer above

A general-purpose AI model, under the Act, is one trained with more than 10²³ FLOPs and capable of producing language, image, or video output. Most commercial foundation models powering enterprise agents clear that bar without much debate.

Article 53 sets baseline duties for every GPAI provider: technical documentation, EU copyright compliance, summaries of training data, and monitoring and incident-reporting policies (the last one only required for systemic-risk models under Article 55). These have applied since August 2, 2025.

Systemic risk status kicks in, presumptively, once cumulative training compute passes 10²⁵ FLOPs. Training at that scale is understood to be extremely costly, so this tier mostly captures the handful of labs operating at frontier scale. Systemic-risk providers face extra duties: model evaluations, incident tracking, and cybersecurity protections covering both the model and the physical infrastructure around it. The AI Office has been explicit that autonomy and tool-use are decisive factors here too, under Article 51(1)(b) and Annex XIII point (e). That means an agentic configuration, heavy tool use, a lot of autonomous action, can push a model that's otherwise borderline into the systemic-risk category. Wrapping a model in agent scaffolding isn't a neutral packaging choice under this Act. It changes the model's regulatory classification.

The European AI Office published the final GPAI Code of Practice on July 10, 2025. It's voluntary, structured around three chapters, Transparency, Copyright, and Safety and Security, and had 20 signatories as of August 2026. Providers who skip it still have to demonstrate compliance some other way and explain their approach to the AI Office directly. The Code already treats agentic behavior as a live consideration in its systemic-risk measures. It's not waiting for some future revision to catch up.

There's a fine-tuning threshold buried in the GPAI Guidelines that a lot of teams miss. If a company modifies a GPAI model using more than a third of the original training compute, or more than 3⅓ × 10²² FLOP when the original figure isn't known, that company becomes a provider under the Act, not just a deployer. Any enterprise fine-tuning a foundation model for its own agentic workflows needs to actually run this math. Crossing that line means picking up GPAI provider obligations on top of whatever the agent itself triggers.

The six high-risk obligations and where agentic systems make each one harder to satisfy

Six obligation areas apply to high-risk systems, spanning Articles 9 through 15, with a compliance deadline of December 2, 2027 for Annex III systems. Each one gets meaningfully harder to satisfy once the system in question is agentic rather than static.

Article 9, risk management. This has to run continuously, through development and through operation, not as a single sign-off before launch. Agents complicate this because they keep changing shape: new tool connections, new MCP server integrations, configuration drift that happens without anyone filing a change request. A risk register updated once a quarter is stale within days. What's actually needed is a live inventory, every agent, every connection, every API, flagged the moment something new gets added.

Article 10, data governance. Data has to be protected from unauthorized access and poisoning across the whole lifecycle, including at inference time, not just during training. Agents call APIs and pull in external data mid-operation, so inference-time access is just as much in scope as the training set was. An anomalous access pattern or a malformed API response now carries operational significance beyond a security event. It's an Article 10 compliance risk.

Article 11, technical documentation. Providers need a complete interface inventory before the system goes to market. Agents that discover and connect to new tools dynamically make this a moving target: documentation written at deployment time can be wrong within hours. The practical fix is a continuously updated, exportable record of every agent, connection, and endpoint, not a PDF filed away after launch.

Article 12, record-keeping and logging. Systems need automatic logging of anything relevant to spotting risk or tracing what happened. Logs have to be tamper-evident and kept for at least six months, longer under other Union or national law for biometric or law-enforcement systems. A log that captures only the model's final output, without the tool calls, the API requests, the data actually touched, the authentication context behind each action, offers only a superficial record. It's a partial record with a gap exactly where the risk lives. Nannini and colleagues state in their April 2026 paper that high-risk agentic systems that can't trace their own behavioral drift can't meet the Act's essential requirements as written.

Article 13, transparency. Deployers and the people affected by a system's decisions need to understand what it does and why. In a multi-agent pipeline, that gets murky fast: which agent in the chain actually made the call, and on what basis? Transparency across a chain of agents handing tasks to each other has to be built into the architecture from the start. Trying to reconstruct it after the fact, from logs never designed for that purpose, doesn't work.

Article 14, human oversight. High-risk systems need to be built so a person can meaningfully supervise them while they're running, with oversight measures scaled to the system's autonomy and risk level. There has to be a working "stop" procedure, a real shutdown path, not a theoretical one. The Act itself calls out automation bias directly: putting a human in the loop does nothing if that human is positioned to just wave outputs through at speed. Meaningful oversight for an agent means the person reviewing it can see what the agent actually accessed and did, not just the summary conclusion it handed back.

Nannini and colleagues (April 2026) go a step further, proposing a twelve-step compliance architecture along with a regulatory trigger map that connects specific agent actions to the laws that apply to them. Their starting point for any provider is blunt: build an exhaustive inventory of every external action the agent can take, every data flow, every connected system, and everyone that action could affect. Skip that step and nothing downstream holds up.

How risk classification works in practice for agentic deployments

Classification runs use case by use case, never system-wide. The exact same underlying agent can be minimal-risk in one deployment and high-risk in another, depending entirely on what it's doing and who it affects.

The Future Society's June 2025 report lays out the decision logic in two steps. First: is this agent a GPAI model, an AI system, or both? Second: does the use case fall into an Annex III category? Agents most often land in Annex III territory through employment and worker management, access to essential services like credit or social benefits, critical infrastructure, or law enforcement.

An agent that autonomously screens job applications, scores creditworthiness, or takes actions that affect someone's access to a benefit or service in an Annex III sector is almost certainly high-risk. That holds regardless of how the vendor markets it or what tier the sales deck claims.

Multi-purpose agents get the same treatment mentioned earlier: designed to operate across several domains, high-risk is the default assumption unless the provider can document real precautions that justify something lower.

This isn't an abstract risk. A company that classifies an agentic HR screening tool as minimal-risk and skips Annex III controls is exposed the moment December 2, 2027 arrives, and exposed to Article 5 scrutiny right now, well before that date. The practical process comes down to two steps done properly: list out every task the agent can perform and every sector it touches, then map each of those tasks against the Annex III categories one by one. Skipping straight to a category label, "it's a chatbot," "it's an assistant," is exactly the mistake the Act's own structure was built to catch.

Sources

  1. EU AI Act Compliance 2026: What High-risk AI Systems Must Do Now | Salt Security
  2. How AI Agents Are Governed Under the EU AI Act
  3. AI Agents Under EU Law
  4. AI Act Service Desk - Frequently Asked Questions
  5. EU AI Act prohibited practices: complete compliance guide for May 2026 | Openlayer
  6. AI Agents Under EU Law
  7. artificialintelligenceact.eu
  8. ai-act-service-desk.ec.europa.eu

More in MCP Policy and Compliance