AI · automation bias · governance · M&E · oversight

Human in the loop is a design decision

What has to be true for a human review step to catch anything, and why the phrase usually appears in a proposal rather than in a workflow.

The phrase appears in the AI section of almost every proposal now. The tool drafts, a human reviews, nothing goes out unchecked. It reassures the donor, it satisfies the ethics annex, and it costs nothing to write.

Then the system ships, and the review step turns out to be a person clicking approve on a screen at four in the afternoon with eleven more to get through before the deadline. The loop exists on the architecture diagram. Whether it catches anything is a different question, and it is a question about how the review step was designed rather than whether it was mentioned.

The distinction matters because the two things fail differently. A system with no human review fails visibly: somebody eventually notices that nobody looked. A system with a review step that nobody can actually exercise fails invisibly, and it fails with a signature on the output.

What the regulation asks for, and what it does not

The EU AI Act is the most specific instrument on this, and it is worth reading closely because it draws exactly the line this post is about.

Article 14 requires that high-risk AI systems be designed so that they can be effectively overseen by natural persons during the period in which they are in use. The word doing the work is "effectively". A human being present is not the standard. The article then lists what the oversight person must be able to do: understand the system's capacities and limitations well enough to monitor its operation and spot anomalies; remain aware of the tendency to over-rely on the output; correctly interpret that output; decide in any particular situation not to use the system, or to disregard, override or reverse what it produced; and halt the system.

Automation bias gets its own clause. Article 14(4)(b) obliges providers to deliver systems in a way that lets the person assigned to oversight remain aware of the possible tendency to automatically rely or over-rely on the output, and it names this as particularly relevant where a system provides information or recommendations for decisions taken by people. That is an unusual thing to find in a regulation. It is a psychological claim, stated as a design requirement.

What the article does not do is say how much disagreement counts as oversight. It requires the capability and says nothing about the exercise. A recent analysis of the article's implementation puts the operational consequence plainly: a production system where the human in the loop approves 99.8 per cent of decisions is exhibiting automation bias, and a system that does not measure that approval rate cannot detect it. Most deployments have no instrumentation at the oversight surface at all.

The regulation is also, for most readers of this blog, not binding. A logframe drafting tool used by a programme team in Nairobi is not a high-risk system under the Act. But the article is the clearest published statement of what a review step has to contain to be worth anything, and the argument does not depend on the enforcement.

The measured cost of a review step

Automation bias is not a theoretical worry. It has been measured, and the numbers are specific enough to design around.

Kate Goddard and colleagues ran twenty-six GPs through twenty prescribing scenarios each, with simulated decision support that was correct seventy per cent of the time and wrong the other thirty. The results are the ones to hold in mind. Clinicians got 50.4 per cent right before seeing advice and 58.3 per cent right after, so the system improved accuracy in 13.1 per cent of cases. It also caused a reversal from a correct answer to an incorrect one in 5.2 per cent of cases. Net improvement, eight per cent.

That is a good system. It is also a system that introduced a new error class which did not exist before it was installed: clinicians abandoning a correct judgement because a machine disagreed with them. Trust in the specific system, decision confidence and task difficulty all predicted how often somebody switched. Less experienced clinicians switched more.

A more recent study put twenty-eight pathology experts under time pressure estimating tumour cell percentages with AI assistance. Overall performance rose significantly, and the automation bias rate was seven per cent: initially correct evaluations overturned by erroneous advice. Time pressure did not make the bias more frequent, but it made it worse when it happened, with heavier reliance on the system's wrong calls and a corresponding drop in performance.

Two things follow. The first is that a review step is not free: it has its own error rate, and that rate is a property of how the step is designed rather than of the reviewer's diligence. The second is that both studies still showed a net gain. The honest reading is not that human review fails. It is that human review has a cost which has to be designed against, and almost nobody designs against it because almost nobody knows the cost exists.

Four questions that decide whether there is a loop

A design decision has answers. Here are the four that separate a review step from a signature.

At what point does the human see it? Before the model runs, setting the constraints? After it drafts, reviewing a proposal? After it acts, catching consequences? The three are different systems. A reviewer who sees the output after it has already gone into a report is not in the loop, they are on the incident response team.

What are they shown? A finished paragraph invites approval. The same content shown as a claim plus the source it came from plus the specific thing that would falsify it invites checking. This is the largest single lever in the whole design and it is usually decided by whoever built the screen, without anyone noticing a decision was made.

What authority do they have? Article 14 puts this third in its list for a reason: the ability to disregard, override or reverse. If the reviewer can reject an output but the deadline means rejecting it costs them a week and rejecting it twice costs them a conversation with their manager, the authority is nominal. Authority in a workflow is measured by what it costs to exercise, not by what the permissions model allows.

How would anyone know it was working? This is the question nobody asks. A review step generates a rate: proportion approved, proportion modified, proportion rejected. If nobody records that rate, nobody can tell the difference between a reviewer who agrees with the model because the model is good and a reviewer who agrees with the model because agreeing is faster. Both look identical from outside and they are the same shape as the 99.8 per cent case.

A team that can answer all four has made a design decision. A team that can answer none has written a sentence in a proposal.

The version of this that applies to results frameworks

Nothing above is specific to M&E, which is partly the point: the oversight problem in a logframe drafting tool has the same shape as the one in a prescribing system, and the medical literature is thirty years ahead.

But the transposition needs care, because the review step in a results framework has a feature the clinical case does not. A prescribing decision has a correct answer that exists independently of the clinician, and the study can score it. Whether an indicator is right for a programme has no such answer. It depends on whether the data source exists in that district, whether the field team can collect it at the frequency stated, whether the definition matches the one the donor's reviewer will apply, and whether the target is achievable given a budget that was cut in the last revision. None of that is in the proposal the model read.

So the reviewer is not checking the output against a known answer. They are checking it against knowledge that exists only in their head, and the design question becomes: does the review surface ask for that knowledge, or does it ask whether the text looks right?

"Approve this indicator" asks the second. "How will data for this indicator reach the reporting system, and who collects it?" asks the first. The second question cannot be answered by anyone who does not hold the local knowledge, which means it cannot be rubber-stamped, which is the property you want.

The cost is that it is slower, and it will be resisted for that reason. A reviewer who has to type an answer for each indicator will process fewer indicators per hour than one who clicks approve, and on a deadline that difference is the whole argument. There is no version of this that is both meaningful and free.

Where this stops

Two limits, and the second one is uncomfortable.

The first is that instrumenting the approval rate tells you something is wrong without telling you what. A ninety-eight per cent approval rate is consistent with a very good model and with a reviewer who has stopped reading. Distinguishing the two requires sampling the approved outputs and checking them independently, which is a second review step, which has its own approval rate. The regress is real and the practical answer is to sample rather than to solve, which is unsatisfying and is what every audit function does.

The second is that the studies above measured expert reviewers with domain training, working on tasks they were qualified for, and still found reversal rates of five to seven per cent. The programme officer reviewing an AI-drafted logframe at four in the afternoon has less training in the failure modes of the system than a pathologist has in the failure modes of a diagnostic tool, and more deadline pressure. There is no reason to expect a better number and some reason to expect a worse one.

What nobody has measured is what that number actually is for AI-assisted results framework drafting. The clinical literature exists because clinical decisions are recorded, scored and audited. Logframe indicators are not scored, so the equivalent study cannot currently be run, and the honest position is that we are designing oversight for an error rate whose size we are guessing at from an adjacent field.


Sources

The regulation

Measured automation bias

Implementation analysis