High-risk AI
AI Act Article 14: human oversight for AI agents
Article 14 of the EU AI Act requires a high-risk AI system to be built so that people can oversee it while it is in use: understand it, override it and stop it. For an AI agent, oversight is a hand-off or an approval step in its tools, not a sentence in its reply.
The article applies to high-risk systems only: from 2 December 2027 for those listed in Annex III, from 2 August 2028 for those under Annex I. This guide covers what it says, who carries it, what it means for an agent and how to test it. It is not legal advice.
What Article 14 says
Paragraph 1 sets the duty. Oversight is designed in, not added afterwards.
High-risk AI systems shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use. (Article 14(1))
Paragraph 2 sets the aim: prevent or minimise risks, including when the system is misused in ways that can be foreseen.
Human oversight shall aim to prevent or minimise the risks to health, safety or fundamental rights that may emerge when a high-risk AI system is used in accordance with its intended purpose or under conditions of reasonably foreseeable misuse... (Article 14(2))
Paragraph 3 makes the measures proportionate to the risks, the level of autonomy and the context of use. They are either built into the system by the provider, or identified by the provider for the deployer to put in place. An agent that acts on its own needs more oversight than one that only drafts.
Paragraph 4 lists what the people assigned to oversight must be able to do.
For the purpose of implementing paragraphs 1, 2 and 3, the high-risk AI system shall be provided to the deployer in such a way that natural persons to whom human oversight is assigned are enabled, as appropriate and proportionate: (a) to properly understand the relevant capacities and limitations of the high-risk AI system and be able to duly monitor its operation... (Article 14(4))
Points (b) and (c) ask that they stay aware of automation bias, the habit of trusting the output too much, and read the output correctly. Points (d) and (e) give them the last word.
(d) to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output of the high-risk AI system; (e) to intervene in the operation of the high-risk AI system or interrupt the system through a ‘stop’ button or a similar procedure that allows the system to come to a halt in a safe state. (Article 14(4))
Paragraph 5 adds a check by two people for remote biometric identification under point 1(a) of Annex III. It rarely concerns a bank's or an insurer's agent.
Who carries it
The provider designs oversight in, and describes the measures in the instructions for use. The deployer puts people in charge of it.
Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support. (Article 26(2))
Article 26(3) leaves the deployer free to organise its own resources to apply the measures the provider indicated. A bank that builds its own agent on a general-purpose model and runs it under its own name is usually both provider and deployer. The deployer guide sets out where the line falls.
What oversight means for an AI agent
An agent acts through tools. It can pre-approve a loan, send a payment or write to a customer. Oversight has to sit where the action happens:
- Approval gates. Above a set amount, the action waits for a person before the tool runs.
- Referral tools. The agent hands the case to a person through a tool call, with the file attached.
- Stop commands. A person can halt the agent, and it stops in a safe state, with nothing left half done.
- Approval through a separate channel. The approval comes through a channel the agent cannot write to, so it cannot approve itself.
- Signals against automation bias. The overseer sees the agent's limits and doubts, not only a recommendation to click through.
This is why an agent that says it escalates is not overseen. "I have passed your request to an adviser" is a sentence. If the referral tool was never called, no adviser will see the file. The control framework for agentic AI turns these measures into controls.
How to test oversight
Test the behaviour, read the tool calls, and send each test ten times. An agent that refers nine times in ten still lets one case through without a person.
Oversight measures and their tests
Some measures can only be proven in documents: who the overseers are, their training and authority, the stop procedure itself. Points (b) to (e) of Article 14(4) are mostly about what the overseer is given, so the evidence is the instructions for use, training records and written procedures. How to test an AI agent explains the split between the agent and the documents.
What Vidimus tests, and what it checks in documents
Vidimus tests the oversight behaviour on the live agent: whether it hands a case to a person when it should, including under pressure. Each test is sent ten times, and the tool calls are read from the reply, the tool log or the event trace. A referral counts only when the referral tool was called. The result is a pass rate with published thresholds; how we test gives them.
A missed referral is filed under Article 14(4)(a) in the evidence pack. The provider's design duties and points (b) to (e) are checked against your documents, passage by passage. Each obligation of Article 14 is checked the way it can be proven, and the pack says which method applied to each.
Until 2 December 2027, these duties do not apply to Annex III systems yet. You can test now: the evidence pack files the results under "Not yet in force", apart from the duties that already apply. The specimen evidence pack shows it: its fictional credit agent missed the referral of an €18,000 application, filed under Article 14(4)(a) with the duties that apply from 2 December 2027.
The evidence to keep
For the wider picture of the high-risk regime, see what applies to high-risk AI systems and the Annex III guide.
To find out whether your agent is high-risk, and so whether Article 14 applies to it, run the free AI Act risk classifier.
Quick answers
Does Article 14 apply to our chatbot?
Only if the system is high-risk: listed in Annex III, or a safety component of a product covered by Annex I. A customer-service chatbot usually is not; an assistant that assesses individuals’ creditworthiness is. Article 50 applies either way.
Is a human in the loop enough?
Only if the agent actually routes cases to that person. A person who is never sent the case oversees nothing. Test that the hand-off happens, through the tool, every time it should.
Who is the overseer?
The deployer assigns oversight to people with the competence, training and authority it needs, and gives them support (Article 26(2)). The provider must make the system possible to oversee.
Can oversight be tested before December 2027?
Yes. For Annex III systems the duties apply from 2 December 2027, but the behaviour can be tested today. Vidimus files those results under "Not yet in force", apart from the duties that already apply.
What does Vidimus test?
The oversight behaviour on the live agent: whether it hands a case to a person when it should, read from its tool calls, ten attempts per test. Design and assignment duties are checked against your documents.
Sources
Vidimus tests your agent’s hand-offs on the live system, ten attempts each, and files the result in a signed evidence pack.