Human-in-the-loop, honestly
Where the loop actually catches errors — and where it just makes the deck look responsible.
- Jurisdictions
- European Union (EU AI Act Article 14), United States (NIST AI RMF; automation-bias laboratory studies), and the WAEMU jurisdictions where the deployments run, which have no equivalent statute
- Evidence period
- 1999-2026
Somewhere in most AI deployment approvals sits one sentence doing enormous work: a human reviews every decision. It appears in vendor contracts, risk registers, board minutes and regulatory filings, converting uncomfortable authorisations into comfortable ones. The decision is whether to accept that sentence as evidence of a control, or to treat it as an unproven claim about a person the board has never met, doing a job it has never costed.
Our position is blunt: a human checkpoint functions as a control only when the human holds four things at the moment of refusal — authority to say no without career consequence, time to actually look, context to see what the machine saw and what it did not, and an incentive structure that does not punish refusal. Remove any one and the checkpoint becomes a costume: it changes who is blamed after the failure without changing its probability. We call that liability theatre; it passes audits until the day it meets a loss.
Start with the law, which is more honest than most deployment plans. Article 14 of the EU AI Act does not say "put a human in the loop." It requires that high-risk systems be designed and developed so that natural persons can effectively oversee them, and enumerates what the overseer must be enabled to do.
Article 14 of the EU AI Act requires that high-risk AI systems be designed so they can be effectively overseen by natural persons, including the overseer's ability to remain aware of automation bias, to interpret output correctly, to decide not to use the system, and to intervene or stop it; measures must be commensurate with risk, autonomy and context.
Article 14 is a list of the conditions under which a human can refuse a machine: the legislator does not assume the human will function, it orders the system built so that functioning is possible.
Automation bias is one of the most stable findings in human-factors science. Skitka, Mosier and colleagues showed in 1999 two characteristic errors: omission, missing events the aid failed to flag even when other indicators showed them plainly, and commission, following the aid against training and fully valid contrary information. Parasuraman and Manzey's 2010 integration added that complacency emerges under multiple-task load — the normal condition of a production review queue — appears in experts as well as novices, and is not eliminated by practice.
Peer-reviewed studies find that humans working with automated aids commit systematic omission and commission errors, and that automation complacency arises under multi-task load, affects experts as well as novices, and is not overcome by practice alone.
NIST's AI Risk Management Framework — voluntary, but the reference vocabulary of most AI governance programmes — places oversight under GOVERN as an organisational design problem and treats the insertion of a human as itself a risk decision: oversight roles inherit the biases the framework catalogues.
The law demands that refusal be possible; the psychology shows it is unlikely under load unless deliberately protected; the framework makes that protection a design task, not a default. What none supplies — because it is firm judgment, not measurement — is our operational test: authority, time, context, incentive, all four, held by the named person on the worst day of the quarter. One consequence is measurable: a checkpoint that has never refused the machine is not evidence that the machine is good, but evidence that the checkpoint is absent. Override rates are a control metric, and zero is a finding.
Evidence cards
CLM-AIACT-ART14-OVERSIGHT-CAPABILITYEU AI Act Article 14 requires high-risk AI systems to be designed for effective human oversight, enumerating the overseer's capabilities (awareness of automation bias, correct interpretation, the option not to use, intervention and stop), with measures commensurate with risk, autonomy and context.
- Context
- Regulation (EU) 2024/1689, text verified 2026.
- Method
- statutory text.
- Contradictory evidence
- the article imposes design capability, not observed effectiveness; the Digital Omnibus on AI (Regulation (EU) 2026/1744, in force 27 July 2026) postponed application of the high-risk obligations.
- Causal confidence
- none claimed — legal fact.
- Transferability
- EU-regulated deployments; persuasive elsewhere.
- Review date
- 2026-08-02.
CLM-AUTOMATION-BIAS-PERSISTENCEhumans working with automated aids commit systematic omission and commission errors; complacency arises under multi-task load, affects experts, and is not eliminated by practice.
- Context
- Skitka, Mosier & Burdick 1999; Parasuraman & Manzey 2010.
- Method
- controlled experiments and integrative review.
- Contradictory evidence
- most findings come from laboratory and simulator settings; magnitude in any specific production deployment must be measured.
- Causal confidence
- experimental for the core effects, in-setting.
- Transferability
- strongest for high-volume review queues under load.
- Review date
- 2026-08-02.
CLM-NIST-RMF-OVERSIGHT-RISKthe NIST AI Risk Management Framework (AI 100-1, January 2023) treats human oversight as an organisational design task under GOVERN and identifies human-AI configurations, including oversight roles, as carrying their own risks.
- Context
- voluntary US framework, current edition 2026.
- Method
- framework text.
- Contradictory evidence
- the framework is non-binding and does not prescribe thresholds or staffing ratios.
- Causal confidence
- none claimed — institutional fact.
- Transferability
- any organisation adopting the RMF vocabulary.
- Review date
- 2026-08-02.
CLM-FIRM-OVERSIGHT-FOUR-CONDITIONSa human checkpoint functions as a control only when the named person holds authority, time, context and incentive to refuse the machine; absent any one, the checkpoint is liability theatre, and a zero override rate is a finding of absence, not of quality.
- Context
- firm operating practice in the corridor.
- Method
- interpretation, not measurement; no deployment statistics cited.
- Contradictory evidence
- in genuinely low-stakes, high-accuracy tasks a lightweight checkpoint may be proportionate; the four conditions are conjunctive by design and unvalidated as a scored instrument.
- Causal confidence
- none claimed.
- Transferability
- bounded — strongest for high-risk decisions reviewed under production load.
- Review date
- 2026-08-02.
The four conditions fail predictably where our clients operate. AI deployments arrive priced on vendor-country staffing assumptions: a queue sized for a European back office, minutes-per-item ratios validated in the vendor's home market. An Abidjan operating company inherits the queue without the staffing plan behind it, and oversight lands on whoever is available, at the end of a reporting line whose real authority sits in Paris or Geneva. Each condition then degrades independently. Authority: the reviewer who refuses is overruling a system the holding company paid for, from the bottom of the organisation chart. Time: the queue arithmetic was never recomputed for local staffing. Context: the model was trained on markets with formal credit histories the corridor does not have. Incentive: throughput is measured, refusal is friction. There is also a regulatory asymmetry: the EU AI Act reaches the European holding company; no equivalent statute operates in the WAEMU jurisdictions where the system runs. Whatever oversight exists in Abidjan exists because the group built it.
Sources and limitations
Sources and limitations. The legal and framework facts rest on the text of Regulation (EU) 2024/1689 Article 14, its application timeline as amended by Regulation (EU) 2026/1744, and NIST AI 100-1; the behavioural findings rest on the peer-reviewed automation-bias literature cited above, whose laboratory provenance is flagged, not hidden. The four-conditions test is the firm's position, graded as interpretation: it is not validated as a scored instrument, and we publish no client oversight metrics because none has passed our evidence-release gate. The transferability boundary is explicit: strongest for high-risk, high-volume decisions reviewed under load; weakest for low-stakes tasks where a light checkpoint may be proportionate. This note is not valid as legal advice on AI Act compliance, nor as a claim that human oversight is useless — only that it must be built, resourced and measured before it is counted as a control.
1. The chief risk officer: publish the override rate of every checkpoint counted as a control. A rate of zero is not comfort; it is the strongest available signal that the checkpoint is a costume.
2. The operations director: recompute the queue arithmetic — items per reviewer per day against minutes of genuine review per item — on the staffing that exists, not the vendor's deck. If it does not close, the time condition has already failed.
3. The head of internal audit: trace one real refusal end to end. Who was questioned, whose metrics suffered, what changed. If none has ever occurred, return to action one.
4. The deployment sponsor: name the checkpoints that exist solely so that a person is blamable, and decide by month-end to remove or resource them. An honest answer names at least one.
Reviewed and countersigned inside the firm before publication: the publication assurer is not the author, and evidence review and French editing sit with a second principal. This is internal role separation, not external or independent peer review.
STG-PUB-NOTE-HUMAN-IN-THE-LOOP
Practitioner observation — not a measured study. No baseline and no sample size are published for this note, so it must not be read as a quantified claim.
- Owner
- Bruno Hounkpati · Operating Chair
- Attribution
- Named public sources cited on the page, each carrying its own evidence grade. Reviewed by Bruno Hounkpati; publication assured by Kevin Abel, Managing Partner.
- Jurisdictions
- European Union, United States, WAEMU member states
- Measurement window
- 1 January 1999 – 31 December 2026
- Baseline
- Not published
- Sample size
- Not published
- Method
- Documentary review of the published sources named on the page. No controlled sample was drawn and no baseline was measured, so this note states an argument from cited evidence, not a quantity of our own.