Oversight is a property of how the system is built, not of a sentence in a policy saying a human reviews the output.
Why it matters when the plan changes
A person nominally in the loop who cannot see why an output was produced, cannot tell how confident it is, and has no practical route to disagree is not providing oversight. They are providing a signature. Article 14 of the EU AI Act makes this a design obligation: high-risk systems must include interface tools that let a person oversee them while in use, not a procedure added afterward. Making oversight real means the output has to carry its reasoning, its confidence and what argues against it.
The tension is that good oversight requires accepting overrides, and overrides make aggregate performance harder to defend. Research on human-AI combinations found the pairing can perform significantly worse than either the best human or model alone, which is why the design carries the result, not the presence of a person. Reliance research points the same way: people abandon a better system after one visible error unless they can adjust its output, in which case both adoption and accuracy tend to rise.
In practice
A recommendation is presented to a leader as a conclusion with a confidence figure and no visible reasoning. They accept it, because disagreeing would mean disputing a number they cannot inspect. The same recommendation with its evidence, its counter-evidence and its blind spots attached produces a different conversation and sometimes a different decision.
Evidence
High-risk systems must be designed so that they can be effectively overseen by people during the period in which they are in use.
Article 14, European Union Artificial Intelligence Act (2024)Combining people and models is not automatically better than either alone, which is why the design of the combination carries the result.
Vaccaro, Almaatouq and Malone, When Combinations of Humans and AI Are Useful (2024)
What it cannot tell you
Human oversight tells you whether a system is built so a person can understand, judge and override its output; it does not tell you whether that person will exercise the override, has the standing to disagree, or possesses the expertise to judge the reasoning shown. Oversight can be present in design and absent in practice.
Questions
That the person can understand the output, judge how much to rely on it, and disregard or override it in practice. Each of those depends on what the system shows them. A review step attached to an output that cannot be inspected produces a signature rather than supervision.
No. Vaccaro, Almaatouq and Malone (2024) found that human-AI combinations performed significantly worse than the best of humans or AI alone. The combination wins only when the person contributes context the model cannot see and the process for combining the two is disciplined, which are design conditions rather than defaults.
Because people abandon a better-performing system entirely after one visible error unless they can adjust its output. Allowing bounded override tends to raise both adoption and accuracy, which makes it a feature of the design rather than a concession to reluctance.
Article 14 of the European Union Artificial Intelligence Act (2024) requires that high-risk systems, including their interfaces, be designed so people can effectively oversee them while in use, with oversight aimed at preventing or minimising risks to health, safety and fundamental rights. The obligation attaches to the design, not to an operating procedure.
The person who owns the consequence of the decision, supported by whoever validated the output. Placing oversight with someone who does not carry the consequence produces the form without the substance, because they have no particular reason to spend effort disagreeing.