Key takeaways in 3 minutes
Putting a person before an Approve button does not create meaningful oversight.
The reviewer needs leverage:
- evidence they can inspect;
- time and attention proportionate to the consequence;
- authority to decide;
- actions beyond approve or reject;
- a safe way to pause, edit, escalate, reverse or stop;
- a record showing what they actually controlled.
Design approval from consequence, reversibility and exposure—not from model confidence or the convenience of the happy path.
There is a small phrase doing heroic work in AI product meetings:
“There will be a human in the loop.”
Splendid.
Which human? In which loop? With what evidence, authority and time? Can they change the proposal, or merely accept the machine's preferred answer? What happens when they press Stop?
If those questions do not have product answers, the human is not a control.
They are a person standing near the control.
If a person carries responsibility but cannot inspect, alter, pause or stop the AI action, they are not in control. They are in the blame path.
Approval Is One Interaction. Oversight Is A Capability.
An approval interface usually appears reassuringly concrete.
There is a named reviewer, a timestamp and a large button. The organisation can point to the screen and say that nothing happens without human approval.
But the design may have already framed the decision:
- the AI recommendation appears first and largest;
- the confidence score implies a correct path;
- the evidence is collapsed;
- Approve is primary while every other action feels exceptional;
- forty unrelated proposals share the same visual weight;
- the reviewer cannot change scope or request more evidence;
- rejecting creates more work than accepting.
The person still clicks.
The interface simply makes one click much easier than considered disagreement.
Research on automation bias has long shown that workload, presentation and reliance on automation can reduce vigilant information seeking. Expertise does not make people magically immune. The practical design response is not to mistrust users. It is to stop designing a conveyor belt and calling the person at the end “oversight”. See Parasuraman and Riley.
Four Conditions For Human Leverage
The Human Leverage Model

1. Knowledge
The reviewer must be able to understand the proposal, its evidence, uncertainty, consequence and system limitations.
This does not mean displaying raw model internals. It means providing decision-relevant evidence at the right level, with deeper inspection available when required.
2. Capacity
The reviewer needs enough time, attention and domain competence to exercise judgement.
A qualified person facing eighty visually identical cards before lunch does not become meaningful oversight through determination alone. Queue design, prioritisation, batching, escalation and interruption frequency are part of the control system.
3. Authority
The person must have the legitimate power to make this decision.
Approval limits, role, jurisdiction, segregation of duties and policy thresholds should be represented and enforced by the product. Routing a £90,000 exception to somebody authorised for £10,000 does not create accountability. It creates a very well-timestamped mistake.
4. Intervention
The person needs actions that match the real work.
Approve and Reject are rarely enough. A reviewer may need to:
- edit the proposal;
- reduce its scope;
- hold it pending evidence;
- request another opinion;
- delegate or escalate;
- choose a safer alternative;
- reverse a prior decision;
- interrupt the system into a safe state.
The EU AI Act's Article 14 uses a similarly concrete vocabulary for human oversight of high-risk AI systems: understand capabilities and limitations, remain alert to automation bias, interpret outputs, disregard or override them, and interrupt the system safely.
Put The Gate Where Consequence Changes
Approval should not be triggered by every model step.
That creates fatigue, teaches people to clear queues and turns genuine control into ceremony.
Classify actions using three dimensions:
- Consequence: what harm, cost or commitment could result?
- Reversibility: how easily and how quickly can the action be undone?
- Exposure: does it affect only an internal draft, or another person, system, customer or market?
Then design a permission ladder:
| Action | Example | Control pattern |
| --- | --- | --- |
| Read-only, low consequence | Inspect delivery history | Run within declared scope. |
| Reversible, contained | Draft a recovery plan | Run, notify and allow edit/undo. |
| Material or externally visible | Change an order date | Preview and require explicit confirmation. |
| Irreversible, regulated or financial | Release payment or delete a legal record | Strong approval, second authority or prohibit. |
Confidence may inform the brief. It should not decide the permission level by itself.
A highly confident system can still perform the wrong high-consequence action. A low-confidence draft may be harmless.
What A Useful Approval Surface Answers
Before a person authorises an action, the interface should answer:
- What exactly will happen?
- To which target?
- Based on which evidence?
- Under which policy and authority?
- What will change outside this screen?
- Can it be undone, and for how long?
- What happens if I reject, hold or escalate?
- Has anything already happened?
The approval should bind to the exact action payload. If the target, amount or scope can change after approval, the person did not authorise the action that eventually executed.
This is where deterministic architecture supports the UX. The model may propose. Identity, permissions, schemas, policy and execution controls should enforce.
Design Friction As A Budget
Not all friction is failure.
Some friction protects attention and makes commitment visible. The design question is where it earns its cost.
Spend friction on:
- high consequence;
- weak or conflicting evidence;
- irreversible action;
- unusual scope;
- policy exception;
- repeated failures;
- changed rules or destinations.
Remove friction from:
- inspecting evidence;
- comparing alternatives;
- correcting the proposal;
- asking for help;
- escalating responsibly;
- undoing a reversible action.
That is a more useful design target than “reduce clicks”.
The Practical Move: Audit The Leverage, Not The Loop
Take one approval screen and score four questions from zero to two:
- Knowledge: Can the reviewer inspect enough to challenge the proposal?
- Capacity: Does the workflow protect time and attention for consequential items?
- Authority: Is the right to decide explicit and enforced?
- Intervention: Can the person meaningfully change what happens next?
Zero means absent. One means present but weak. Two means usable and enforced.
Do not total the numbers and congratulate the average. Find the zero.
That is where the “human in the loop” is most likely to be decorative.
The Accountability Lab Connection
One strand of Accountability Lab is the approval-pattern gallery: making the difference between nominal approval and meaningful control visible enough to critique.
That includes the unglamorous details—authority tiers, equal-weight actions, context engagement, escalation, safe interruption and the record of what the person could actually do.
Because “a human was present” is not a design outcome.
“A human had meaningful leverage” is much closer.
Sources And Further Reading
- EU AI Act, Article 14: Human oversight — requirements apply to high-risk AI systems.
- ICO guidance on meaningful human oversight.
- Parasuraman and Riley, Humans and Automation.
- Calvert et al., Meaningful Human Control.
- OWASP LLM06:2025 Excessive Agency.



