On average, human-AI combinations performed worse than the better of humans or AI alone, with losses in decision tasks and gains in content creation.
Why it matters when the plan changes
Putting a human in the loop is often treated as a safeguard that also improves results. The evidence is less comfortable. Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone's 2024 meta-analysis of 106 experimental studies, reporting 370 effect sizes, found that combinations on average did worse than the best of humans or AI alone, with performance losses in tasks that involved making decisions.
The tension is that combinations can still work. The same analysis found significantly greater gains in tasks that involved creating content, and other research shows people use models more when they can adjust them. Complementarity depends on defining what each side contributes, such as the model's consistency and the person's context and authority. Adding a human without that definition adds cost and can subtract accuracy. The combination has to be tested against both parts, not assumed.
In practice
A team reviews an AI-assisted process for flagging contract risks. Lawyers see every flag and approve or reject it. An audit shows the lawyers overturn correct flags almost as often as wrong ones. The redesign has the model flag and explain, lawyers add facts the model cannot see, and both the model's call and the final call are recorded and compared.
Evidence
A meta-analysis of 106 experimental studies reporting 370 effect sizes found that, on average, human-AI combinations performed worse than the best of humans or AI alone.
Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone, When Combinations of Humans and AI Are Useful, Nature Human Behaviour (2024)The same analysis found performance losses in tasks that involved making decisions.
Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone, When Combinations of Humans and AI Are Useful, Nature Human Behaviour (2024)People were considerably more likely to use an imperfect algorithm when they could modify its forecasts.
Berkeley J. Dietvorst, Joseph P. Simmons and Cade Massey, Overcoming Algorithm Aversion, Management Science (2018)
What it cannot tell you
Complementarity is measured per task and setting, and most evidence comes from experiments rather than live organisational decisions. A combination that works in one task can fail in another, and the meta-analysis averages across very different designs, so its headline finding is a warning rather than a verdict on any specific process.
Questions
Not automatically. A 2024 meta-analysis in Nature Human Behaviour of 106 experimental studies found that, on average, human-AI combinations performed worse than the better of humans or AI alone. Gains appeared in content creation; losses appeared in tasks that involved making decisions.
One plausible reason is that people adjust the model where it was right and defer to it where it was wrong, so errors from both sides survive. When the person has no information the model lacks, their adjustments add noise rather than signal, and the combination drifts below the better of the two.
Defining each role. The person should contribute context the model cannot see and be able to understand, question, alter or reject its output. Berkeley Dietvorst and colleagues found in 2018 that people used an imperfect algorithm considerably more when they could modify its forecasts.
Oversight is about control and accountability: someone able to understand, intervene and stop a system. Complementarity is about performance: whether the combination beats either part. A process can satisfy oversight rules and still lower accuracy, so the two have to be assessed separately.
By measuring three things on the same cases: the model alone, the person alone and the combination, against outcomes defined in advance. Only then can the combined process be shown to add value. Testing the model alone, or asking reviewers whether they found it helpful, cannot answer the question.