KodDeltaGuides

Human-in-the-loop AI: designing the approval step

~11 min read

Human-in-the-loop means the AI prepares the work and a person makes the decision. Where the approval sits is settled by one test: is the action reversible? Every irreversible step — payment, shipment, contractual commitment, permanent deletion — needs a person on the button. If an approval takes longer than about thirty seconds, nobody will do it properly.

The most expensive mistake in AI projects is not technical. It is authority design. Either the system performs an action nobody authorised and the damage surfaces late, or approvals are placed everywhere, the queue backs up, nobody clears it, and the system is abandoned. Both failures come from the same missing decision.

This article covers where the approval step belongs, how to design the screen, what to record, and what EU, Turkish and Gulf regulation now expects.

One test: is it reversible?

The position of every approval step is settled by a single question: if this action is wrong, can it be undone, who notices, and how long does that take?

Three answers, three behaviours.

Reversible and it leaves a trace. No approval needed. Creating a record, extracting data, tagging, drafting, classifying, searching. If it is wrong, it gets corrected and the correction is logged.

Reversible but silent, or noticed late. Sample-based audit. The system runs; a defined percentage of its output is checked by a person on a regular cadence.

Irreversible. Approval is mandatory. Cash movement, stock allocation and shipment, an outbound commitment to a customer, contract execution, permanent deletion, and any write to an external system that cannot be retracted.

This three-way split makes the “how reliable is AI” debate unnecessary. A system that is right 99% of the time is still not left alone on an irreversible step, because the remaining 1% cannot be undone.

Three oversight models and where each belongs

ModelHow it worksFitsIts failure mode
Human-in-the-loopEvery decision is authorised; nothing executes unapprovedPayment, shipment, contracts, outbound messagesQueue backs up, rubber-stamping begins
Human-on-the-loopSystem runs, a person monitors and intervenesClassification, routing, extractionMonitoring stops happening in practice
Human-in-commandAuthority to stop and disable stays with a personMinimum baseline for every systemThe stop button has never been tested

Article 14 of the EU AI Act requires human oversight for high-risk systems and specifies that the person overseeing must be able to understand the system’s capacities and limitations, interpret its output correctly, disregard or override that output, and safely stop the system. The same article explicitly names awareness of automation bias — the tendency to accept output without checking it.

The practical point is that you use all three models in one system. In an invoice matching build, extraction and matching are human-on-the-loop, payment is human-in-the-loop, and the whole system is human-in-command.

A worked example: the approval screen before payment

The process. Checking supplier invoices and preparing them for payment.

What the system does. Reads the invoice, matches it against the purchase order and goods receipt, writes a reason for every out-of-tolerance difference, prepares a payment proposal.

What is on the approval screen.

  • The original invoice image on the left, the extracted lines on the right.
  • Match status and confidence beside every line.
  • A plain-language reason on out-of-tolerance lines: “unit price 4% above the order”.
  • Links to the purchase order and goods receipt records.
  • One approve button and one send-to-exception button.

What is deliberately absent. A bulk approve control. Approving forty invoices with one click turns the approval step into a ceremony.

What is recorded. Who approved, when, what was on their screen, what the system proposed, and what the person changed. Those five fields satisfy the audit requirement and simultaneously form the only dataset that shows where the system gets things wrong.

How you know it is working. Two indicators. If average approval time collapses over the first months, rubber-stamping has started. If the correction rate declines, the system is genuinely improving.

Five rules for approval screen design

1. The thirty-second rule. If an approval averages more than about thirty seconds, people stop doing it properly. The fix is never to push the person harder; it is to choose what appears on the screen more carefully.

2. The source is one click away. If the approver has to open another system to check something, the screen is not doing its job.

3. Confidence must be visible. Items the system is unsure about need to look different from items it is sure about. If everything looks the same, everything gets the same attention — which means none of it gets enough.

4. Rejecting must be as easy as approving. Screens that require a written justification to reject are screens that push people toward approving.

5. No bulk approval. The only exception is a group the system itself assembled under one source, one rule and one reason, with the group contents visible.

What the audit record must contain

The half of approval design most often missing is the record. “Who approved it” is not enough; a year later, someone must be able to reconstruct what the decision was based on.

FieldWhy it is needed
The system's proposalSo it is clear what the person accepted
Evidence and source referencesTo separate a bad document from bad logic
Confidence levelTo measure how low-confidence items are treated
Approver and timestampTo establish the chain of responsibility
What was displayedTo know what information the decision used
What the person changedTo learn where the system is wrong
Time taken to approveTo monitor rubber-stamping

The last row is almost never captured and is the most informative field on the list. The trend in approval time is the only measurable signal that oversight is real rather than nominal.

Retention is a separate decision. If the records contain personal data, keeping them indefinitely conflicts with data minimisation; deleting them destroys the audit trail. The workable compromise is to retain the decision record and mask the content after the retention period.

Catching approval fatigue early

Placing a human approval does not guarantee the decision stays with the human. As the system keeps being right, the approver stops checking. This is not a design defect so much as a known property of human behaviour, and the countermeasures belong in the design.

  • Sample audit. A percentage of approved items is reviewed by a second person. The purpose is not to catch individuals; it is to catch the system degrading quietly.
  • Approval time monitoring. An alert when the average drops sharply.
  • Difficulty separation. Low-confidence items go into their own visually distinct queue.
  • Rotation. The same person looking at the same queue for months loses acuity.
  • Share the system’s mistakes. An approver who knows where the system tends to fail pays attention in the right places.

Four warning signs, any one of which means the design has broken: approval time in single-digit seconds, a correction rate near zero while exceptions are still being produced, the queue being cleared in one batch at the end of the day, and a single approver whose absence stops the process.

All four have the same remedy — send less to the approval step. Widening tolerance rules, moving low-risk items to sample audit, and reviewing auto-passed items in a weekly report together keep the queue survivable.

The regulatory picture

European Union. Article 14 mandates human oversight for high-risk systems. The Act’s transparency obligations became applicable on 2 August 2026, requiring disclosure where a person is interacting with an AI system. Some high-risk obligations were deferred; the transparency article was not.

Türkiye. The data protection authority published guidance on generative AI in November 2025 and on agentic AI systems in March 2026, emphasising data minimisation, purpose limitation and the reviewability of automated decisions.

Gulf. Saudi Arabia’s personal data legislation and the AI adoption framework published by the national data and AI authority make data governance, model accountability, transparency and human oversight explicit headings.

Three regimes, one shared requirement: a person who is competent, informed and genuinely able to intervene. That is not satisfied by adding an approve button late in the project. It has to be how the system is built.

When not to add an approval step

Honesty runs both ways here. Unnecessary approvals do as much damage as missing ones.

On reversible steps. Creating records, drafting, tagging, classifying. Approvals here generate a queue, and queues consume attention.

When the approver cannot actually decide. An approval from someone without the information or authority to change the outcome distributes responsibility and nothing else.

When it would generate more than about 200 approvals a day. At that volume approval becomes ritual. Use tolerance rules and sample audit instead.

When the purpose is to spread blame. Approvals added so that “someone has looked at it” become indistinguishable from real approvals and devalue both.

When a clear rule already covers it. Approving a match that fell inside a tolerance you wrote makes writing the tolerance pointless.

Measure your own situation

For each step in the process you plan to automate, answer four questions. The resulting table designs the oversight for you.

  1. Is this step reversible? Yes or no. No means approval is mandatory.
  2. If it is wrong, who notices and how quickly? “Months later” means you need sample auditing.
  3. What does an error cost, per occurrence? Write a figure. It determines how much time the approval step deserves.
  4. Who approves, and how many approvals a day will that be? A name and a number. Above 200, change the design before building.

Written alongside your process steps, those four answers become the best technical specification you can hand a supplier. You will have set the oversight boundaries yourself rather than accepting whatever the vendor’s product happens to do.

Next step

Our approach to operational automation is on the AI process automation page, the organisations we build for on who we build for, and the delivery bands on the pricing page. Concrete approval points are worked through in invoice matching automation and sales order entry automation. Fill in the four-question table above and send it through the quote form; we will design the oversight around it.

Frequently asked questions

Doesn't human approval cancel out the savings?

No, because the saving comes from preparing the decision, not from making it. Reading an invoice and comparing it to an order takes minutes; confirming the result takes seconds. When the approval step is well designed it is a small share of total time. What destroys the saving is a badly designed approval screen, not the approval itself.

What are the three oversight models?

Human-in-the-loop: a person authorises each decision. Human-on-the-loop: the system runs and a person monitors and intervenes. Human-in-command: a person retains the authority to stop or disable the system. All three can coexist in one system, with different models on different steps according to risk.

What should the approval screen show?

Four things: what the system proposes, what evidence it used, a one-click link to the source, and how confident it is. If any of the four is missing, the approver either rubber-stamps or re-does the work from scratch — and both outcomes make the automation pointless.

What is automation bias?

The tendency to accept a system's proposal without checking it. It means the decision stays with the machine even though an approval step exists. Countermeasures are design-side: display confidence, separate low-confidence items into their own queue, sample-audit approved items, and monitor approval speed. A sharp drop in approval time means rubber-stamping has begun.

What does regulation require?

Article 14 of the EU AI Act requires human oversight for high-risk systems, and specifies that the overseeing person must understand the system's limitations, interpret its output correctly, be able to disregard or override it, and be able to stop it. Gulf frameworks name human oversight explicitly as well. The shared requirement is a person who is competent, informed and genuinely able to intervene.

Won't approvals everywhere slow the system down?

Yes, which is why you do not put them everywhere. Separate steps by reversibility. Creating a record, drafting, extracting data and classifying are reversible and need no approval. Cash movement, stock allocation, outbound commitments and permanent deletion are not, and do.

Related guides

Service page: Custom software service

Let's talk about what you need.

The 30-minute discovery call is free and carries no commitment.