Blog
Insights
Human in the Loop for AI Agents: Which Actions Should Wait for a Person
Human in the loop for AI agents: which actions should wait for approval, how to avoid approval fatigue, and what EU AI Act Article 14 expects from oversight.

Mario Baburic
Founder & CEO

By Dr. Mario Baburić, Founder & CEO, Booga Enterprise
An agent finishes a supplier review at 18:05, drafts the email that ends the contract, and sends it before anyone in procurement has read it. The analysis may well have been correct. The decision to send it was never taken by a person.
Human in the loop for AI agents is the control that closes that gap. The organisation decides in advance which actions an agent may take on its own and which ones wait for someone with the authority to approve them. Applied well, it puts people at the points where their judgement changes the outcome and keeps them away from steps where review only adds delay.
This article sets out where human approval belongs, why both extremes fail, how to keep approvals meaningful once they are in place, and what the EU AI Act expects from human oversight.
What human in the loop means for AI agents
An AI assistant answers a question and stops. An AI agent plans a task, calls tools and data sources, and produces a result, often across several systems in a single run. The difference between an AI assistant and an AI agent changes where the risk sits. With an assistant, a person reads every output before anything happens. With an agent, the output can be an action: a record written to a CRM, an email sent to a client, a configuration changed in production.
Human in the loop (HITL) places a person between the agent's proposal and its execution for a defined set of actions. The agent does the work up to the point of consequence. A person reviews what it intends to do, with the context needed to judge it, then approves, edits or rejects. The decision and the reviewer are recorded.
Two related terms are worth keeping apart. Human on the loop describes oversight after the fact: people review logs, exceptions and aggregated outcomes while the agent acts independently. Human out of the loop means no review at all. Most agent programmes need all three, applied to different actions.
Why approving everything and approving nothing both fail
In May 2026, Gartner predicted that by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps that only surface after production incidents. Its analysis names the cause: organisations treat AI agent governance as binary, with every agent either locked down or fully trusted.
Each extreme fails in a predictable way. When every agent sits behind the same heavy approval process, simple agents that only summarise or retrieve information slow down, and teams route around the controls. Gartner identifies this as a driver of shadow development, the same pattern behind shadow AI becoming an agent problem inside enterprises. When every agent is trusted equally, agents that write data, send messages or change configurations run under the same light controls as a document summariser, and the first serious error is found in production.
The commercial cost is already in Gartner's earlier forecast: over 40% of agentic AI projects will be canceled by the end of 2027, with inadequate risk controls named alongside escalating costs and unclear business value. Proportional human oversight is one of the most direct ways to address the risk-control part of that list.
Four autonomy levels, and where human approval belongs
Gartner recommends classifying AI agents by autonomy level, with each level carrying its own trust boundary and its own controls. The four levels give a practical map of where human in the loop applies and where a lighter form of oversight is enough.
Level | What the agent can do | Where the person sits | Governance focus |
1. Observe | Read-only access to defined data sources; output visible only to the requesting user | Uses the output | Scoped data access, user authentication, usage logging, functional and security testing |
2. Advise | Drafts, recommendations and proposed actions; no write access to any system | Reviews every output and carries out actions manually | Level 1 controls plus accuracy and hallucination testing, quality evaluation and training on appropriate reliance |
3. Act with approval | Writes data, sends communications or changes configurations, only after explicit approval for each action | Approves each action before it runs | Clear approval workflows with audit trails, security testing and agent-specific incident response |
4. Act autonomously | Executes actions independently within defined guardrails | Reviews exceptions, audit logs and aggregated outcomes | Continuous monitoring, enforced guardrails, rapid rollback, circuit breakers and named ownership |
Source: Gartner, AI agent autonomy levels, May 2026.
In practice, the level attaches to each action an agent can take. One agent can read a contract repository at Level 1, propose a clause change at Level 2 and wait for approval at Level 3 before it emails the counterparty. Classifying the whole agent at its riskiest action over-restricts the safe steps; classifying it at its safest action leaves the risky one unchecked.
How to decide which AI agent actions need human approval
Five questions settle most cases. An action that scores high on any one of them is a candidate for human in the loop.
Reversibility. Can the action be undone fully and at low cost? A draft can be discarded. A sent email, a payment or a deleted record cannot.
Reach. Does the result leave the organisation, or reach a client, counterparty, regulator or the public?
Data sensitivity. Does the action touch personal data, privileged material, financial records or another client's information?
Financial and contractual weight. Does it commit money, change a price or alter terms?
Track record. Is this a new agent, a new workflow or a new data source? Early runs justify an approval step that can be relaxed once the run history supports it.
Applied to common agent tasks, the questions produce a placement like this:
Agent action | Starting level | Why |
Summarise the last ten board minutes for the requesting director | 1. Observe | Read-only, internal, output seen only by the requester |
Draft a reply to a client query from the knowledge base | 2. Advise | Reaches a client, so a person reviews and sends it |
Update a deal stage in the CRM | 3. Act with approval | Writes data; can move to Level 4 with sampled review once the run history is clean |
Send a payment reminder to a customer | 3. Act with approval | External, financial and hard to retract |
Post the weekly market brief to an internal channel | 4. Act autonomously | Internal and reversible; exceptions reviewed |
Change a production configuration | 3. Act with approval | Operational risk; approval by a named owner of that system |
A short rule covers most of the table. If an action is irreversible and leaves the organisation, it waits for a person. If it is reversible and stays internal, it can run on its own, provided every run is recorded and someone reviews the exceptions.
Keeping human approval meaningful: avoiding approval fatigue
An approval step only protects the organisation while the person approving reads what they are approving. Gartner warns that at the act-with-approval level, approvals can degrade under time pressure or approval fatigue and create a false sense of safety. Five design choices keep human in the loop working after the first month:
Show the decision, with its context. The reviewer sees the proposed action, the inputs and sources the agent used, and what will happen on approval.
Route approvals by authority. Each request goes to someone entitled to approve that specific action, which is where capability-based permissions and runtime access control for AI agents come in.
Record every decision. The approver, the time and the exact action approved are written to the audit record, so the approval becomes evidence. The contents of an enterprise AI audit trail should include this control state for every step.
Watch the approval pattern. A queue where nearly every request is approved within seconds suggests the step is either unnecessary or not being read. Either result is worth acting on.
Review placement on a schedule. Move actions between levels on the basis of run history, in both directions.
These choices belong in the platform the agents run on. Approval steps added by convention outside it are the first thing to be skipped under deadline pressure. This is the same argument behind AI contextual governance: controls enforced where the agent executes hold up; controls kept in a policy document tend not to.
Human oversight under the EU AI Act
Article 14 of the EU AI Act requires high-risk AI systems to be designed so that natural persons can oversee them effectively while they are in use. Oversight measures must be commensurate with the risks, the level of autonomy and the context of use, the same proportional principle behind the autonomy levels above. Article 14 also sets out what the people doing the oversight must be able to do, including understanding the system's capabilities and limitations, remaining aware of automation bias, interpreting outputs correctly, deciding to disregard or override an output, and intervening in or stopping the system.
Under the Digital Omnibus on AI, Regulation (EU) 2026/1744, obligations for stand-alone high-risk systems listed in Annex III apply from 2 December 2027, a little over 13 months from the publication of this article. Whether a particular agent falls into a high-risk category depends on its use case, so classification is a question for legal and compliance teams. The design question is the same in every case: can a person see what the agent is about to do, stop it, and show afterwards that they did. Our EU AI Act guide for enterprise AI agent deployments covers the dates and obligations in more detail.
How Booga One and Booga Agents keep a person in the loop
The control questions are the same for one professional running an agent and for an organisation running many. Booga One and Booga Agents answer them at those two scales.
In Booga One, agent runs that need sign-off wait in Pending Approvals on the dashboard until you review them. The Monitor tab in Agents shows every run step by step, with the output of each step, so you can see what the agent did before you rely on the result. Personas and Collections limit which documents and sources the assistant draws on, which keeps an agent working from the material you chose.
Try Booga One free. Free tier, no card required: start building your first agent.
In Booga Agents, the same principle runs at organisational scale. Capability-level RBAC is enforced at runtime, so whether a user can trigger an action is checked when the action runs, against that user's capabilities. User and system activity is captured in the audit pipeline, with 14 event categories and a 7-year default retention that each tenant can configure, and audit events stream to the organisation's SIEM. Booga Agents runs on dedicated resources in the customer's own cloud and region, on Azure, AWS or GCP.
Teams that prototype an agent in Booga One bring its definition and workflows across to a dedicated Booga Agents tenant by export and import when the work needs organisational controls. Prototype in Booga One. Operationalize in Booga Agents.
Where to start
Autonomy is granted one action at a time. Before the first agent reaches production, decide which of its actions run on their own, which wait for a person and who that person is. Record every decision, and move actions between levels when the run history supports it. Organisations that set this up early can extend agent autonomy with evidence behind each step; those that skip it tend to find the gaps through an incident.
Booga Agents is available to companies by request. Request access to Booga Agents to discuss human oversight, audit and deployment on your own cloud.
Frequently asked questions
What does human in the loop mean for AI agents?
Human in the loop for AI agents means a person reviews and approves specific agent actions before they are executed. The agent prepares the action, a person with the right authority approves, edits or rejects it, and the decision is recorded. It is applied to actions with real consequences, such as sending external communications, writing data or committing money.
Which AI agent actions should require human approval?
Approval belongs on actions that are hard to reverse, reach people outside the organisation, touch sensitive data, carry financial or contractual weight, or come from a new agent or workflow. Read-only and internal, reversible actions can usually run without approval, provided they are logged and exceptions are reviewed.
What is the difference between human in the loop and human on the loop?
Human in the loop requires approval before an action runs. Human on the loop lets the agent act within guardrails while people review exceptions, logs and aggregated outcomes afterwards. Mature agent programmes use both, assigned per action according to risk.
What is approval fatigue in AI agent governance?
Approval fatigue is the point at which reviewers approve requests without reading them, usually because there are too many approvals or too little context in each one. Gartner warns it can create a false sense of safety. Approving only consequential actions, showing context with every request and monitoring approval patterns reduce it.
Does the EU AI Act require human oversight of AI agents?
Article 14 of the EU AI Act requires human oversight for high-risk AI systems, with measures proportionate to risk, autonomy and context of use. Under Regulation (EU) 2026/1744, obligations for stand-alone Annex III high-risk systems apply from 2 December 2027. Whether a specific agent is high-risk depends on its use case and should be assessed by legal and compliance teams.

Mario Baburic
Founder & CEO
Share


