In this article9 min read
The paper lands before the meeting. An AI summary sits at the top, followed by a recommendation and a polished explanation of why it is right.
By the time the first person speaks, the difficult part may already have happened. The evidence has been framed, the plausible interpretation has been supplied and disagreement now has to overcome a complete-sounding answer.
That is the practical problem at the centre of AI leadership judgement. The risk is not simply that a model can be wrong. It is that a persuasive rationale can make independent inspection feel unnecessary precisely when it matters most.
A human is not meaningfully in the loop if the machine has already supplied the frame, the recommendation and the justification.
AI leadership judgement needs a better evidence boundary
Leonid Sudakov and Nathan Furr raise this concern in their Harvard Business Review article, AI Is Undermining Leaders’ Judgment. Here’s What to Do About It. They describe original judgement through two capacities: noticing signals outside the dominant view and forming an interpretation that does not merely defer to the model or the group. Their organisational response is to build structured curiosity and intentional dissent.
That is a useful leadership argument. The field experiment behind it is narrower—and more interesting—than the headline suggests.
The underlying Harvard Business School working paper studied 228 people screening 48 real submissions to the 2024 MIT Solve Global Health Equity Challenge. They produced 3,002 pass-or-fail decisions in three conditions:
- no AI recommendation;
- a GPT-4 pass-or-fail recommendation without a rationale;
- the identical recommendation accompanied by a narrative rationale.
Four MIT Solve-affiliated experts provided the comparison benchmark. Two experts assessed each submission and agreed on 42 of the 48 cases. This was structured expert judgement, not an objective record of which innovations later succeeded.
+4.3 points
Bare AI recommendations improved adjusted agreement with the expert benchmark versus the control.
No gain
Adding a fluent rationale produced no significant improvement over the control.
−11.9 points
Narratives reduced useful overrides when the model disagreed with the expert benchmark.
The surprising result was not that AI failed. A bare recommendation improved adjusted agreement with the expert benchmark by 4.3 percentage points. The narrative version did not. Black-box assistance significantly outperformed narrative assistance even though both displayed the same underlying recommendations.
The clearest difference appeared when the model was wrong. When its recommendation conflicted with the expert benchmark, evaluators followed it in 54.7% of black-box cases and 65.3% of narrative cases. The explanation made an incorrect answer 10.6 percentage points harder to resist.
This did not prove that AI generally weakens leaders’ judgement. Only 72 participants had direct MIT Solve connections; the remaining 156 came from university programmes. The task was a time-pressured first-stage screen, not executive strategy, and every condition could open an AI-generated submission summary. The tested model was an April 2024 version of GPT-4.
It did demonstrate something operationally important: in this setting, a fluent explanation increased deference when independent disagreement would have improved the decision.

A polished explanation is not the same as inspectable reasoning
A narrative generated by a language model can be clear, concrete and internally coherent. Those qualities make it useful for communication. They do not prove that the explanation faithfully exposes the mechanism that produced the recommendation.
In the experiment, people who accepted narrative recommendations viewed 8.3 percentage points less submission content and scrolled 8.6 points less than the control group. These are behavioural proxies, not a direct measurement of thought, and the compliant subgroups were not separately randomised. Even with that caveat, the pattern fits a recognisable failure mode: the explanation becomes a substitute for examining the source.
That is why “human in the loop” is an incomplete control. It describes where a person appears in a process, but not:
- whether they saw the evidence before the recommendation;
- whether they recorded an independent view;
- whether they have the authority and time to disagree;
- whether dissent is visible in the decision record;
- whether anyone later checks which overrides were productive.
NIST’s AI Risk Management Framework guidance on human-AI interaction makes the same governance problem visible. Human roles and responsibilities need to be defined, and organisations need evidence about when people challenge system outputs—not simply a final approval box.
The UK Government’s current Data and AI Ethics Framework similarly calls for named oversight, validation and human responsibility in risky or high-impact uses. The field experiment adds a warning: responsibility without workflow design can still leave the reviewer inside the model’s reasoning.
Use AI after the first view, not before it
The most practical safeguard is sequence. For consequential, subjective decisions, require a short independent view before exposing the reviewer to an AI recommendation or rationale.
This is not a demand for long, unaided analysis. A two-minute note can be enough if it captures four things:
- Observation: what the source evidence actually says.
- Interpretation: what the reviewer currently thinks it means.
- Uncertainty: what is missing, ambiguous or outside their expertise.
- Reversal condition: what evidence would change the view.
Only then should AI enter. Its most useful role is not to make the first view sound more complete. It is to challenge that view: surface omitted evidence, offer a competing interpretation, identify assumptions and explain what would have to be true for the opposite decision to be better.
The first view still has to be grounded in reality. Observing frontline work, reading the original material and hearing affected voices all widen perception before a summary compresses it.
This also extends the argument in The 80% Problem. Cognitive offloading becomes dangerous when we stop distinguishing saved effort from surrendered judgement. Recording a first view preserves that distinction without rejecting the efficiency of AI.
Turn dissent into a role, not a personality test
Sudakov and Furr’s second practice—intentional dissent—matters because disagreement rarely survives if it depends on one person feeling brave in the room.
Make challenge part of the operating model:
- Separate proposal from challenge. The person who used AI to prepare the case should not be its only reviewer.
- Assign the minority case. Ask someone to state the strongest evidence-based argument against the emerging recommendation.
- Challenge the frame. Test whether the decision question, threshold or comparison set excludes a legitimate option.
- Protect the trace. Record disagreements and unresolved uncertainty instead of editing them out of the final paper.
- Review the override. When outcomes become visible, ask whether following or rejecting the AI improved the result.
This is different from ritual devil’s advocacy. The objective is not to manufacture opposition. It is to preserve an alternative interpretation long enough to test it.
The same discipline appears in the Decision Monitoring Loop: the quality of a decision depends partly on whether the system can observe its own assumptions, uncertainty and error after action. A dissent record turns that learning into evidence.

Match the control to the cost of being wrong
Not every AI-assisted decision needs the full protocol. Friction should follow consequence.
Use a light check when
- the decision is easily reversible;
- the criteria are objective;
- feedback arrives quickly;
- a mistake has limited impact.
Use the full review when
- the judgement is subjective;
- rejection removes future options;
- feedback is delayed or absent;
- the decision affects people, money, safety or rights.
The error balance matters too. In the MIT Solve experiment, AI assistance reduced false positives: fewer submissions advanced when the expert benchmark would have rejected them. But when the AI recommended rejection, false negatives rose by 8.4 percentage points with a bare recommendation and 14.9 points with a narrative. The explanation improved defensibility while making promising options easier to discard.
Leaders therefore need to decide which error is more dangerous before configuring the workflow. A conservative filter may be appropriate for a reversible compliance check. It may be destructive in innovation, hiring or early discovery, where a rejected option is unlikely to return.
This is also an accountability question. As explored in the problem of AI narratives and responsibility, a plausible system explanation can make an organisational choice appear inevitable. The decision record should show who accepted the recommendation, what evidence they examined and which alternatives they considered.
Download the Independent Judgement Review Pack
The free seven-page pack below turns the argument into a working review method. It is designed for strategy papers, investment choices, innovation screens, supplier decisions and other cases where AI helps interpret evidence but a person remains accountable.
- a consequence and error-cost triage;
- a decision frame with accountable owner and evidence standard;
- a first independent view completed before AI input;
- an AI challenge and source-verification sheet;
- an intentional-dissent record;
- a final decision trace and calibration review.
It does not award a mechanical score or pretend there is one universal approval threshold. Its purpose is to keep observation, interpretation, AI influence, dissent and accountability visible.
What the research can—and cannot—tell us
The study is a 64-page working paper, last revised in February 2026, rather than a final peer-reviewed journal article. Its primary AI-versus-control comparison was pre-registered, but the false-negative, decision-quality and mechanism hypotheses were not pre-registered in their current form. The expert benchmark also replaced the originally planned external innovation ratings.
It tested one model, one explanation format, one threshold and one early-stage screening process. It measured short-session decisions, not long-term loss of skill. Its false negatives were disagreements with four experts, not innovations proven successful later. The result should not be generalised automatically to medicine, audit, recruitment or executive strategy.
A separate Microsoft Research and Carnegie Mellon survey of 319 knowledge workers found that higher confidence in generative AI was associated with less reported critical-thinking effort, while higher confidence in one’s own ability was associated with more. That study was self-reported and cross-sectional, so it does not establish long-term cognitive decline either. Together, the studies justify better workflow design, not a claim that using AI inevitably erodes judgement.
The review protocol and downloadable pack are Beta Tester Life interpretations of this evidence. They have not been experimentally validated and are not substitutes for legal, regulatory, safety, HR or professional advice.
Sources used
- Leonid Sudakov and Nathan Furr, AI Is Undermining Leaders’ Judgment. Here’s What to Do About It., Harvard Business Review, 19 August 2026.
- Jacqueline N. Lane, Léonard Boussioux, Charles Ayoubi, Ying Hao Chen, Camila Lin, Rebecca Spens, Pooja Wagh and Pei-Hsin Wang, The Narrative AI Advantage? A Field Experiment on AI-Augmented Evaluations of Early-Stage Innovations, Harvard Business School Working Paper 25-001, revised 26 February 2026; DOI record.
- Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks and Nicholas Wilson, The Impact of Generative AI on Critical Thinking, CHI 2025.
- National Institute of Standards and Technology, AI Risk Management and Human-AI Interaction, AI RMF 1.0 Appendix C.
- UK Government, Data and AI Ethics Framework, 2025.
AI can widen the evidence available to a leader. It can also make one interpretation feel complete before alternatives have been seen. The scarce capability is not producing another answer. It is preserving enough independent judgement to recognise when a persuasive one deserves to be refused.
Continue the journey
One thought leads to another.
Scroll to explore
AI Governance
AI Build Versus Buy: How to Make Better Ownership Decisions
AI Governance
AI Is Not a University. But It Can Build You One.
Leadership & Culture
The Hidden Value of Fleeting Encounters in a Transactional World
AI Governance
The Hidden Verification Gap: A Practical 30-Day Way Forward
Field Notes
The CC List Is Not a Control
Enterprise Delivery
Stop Fixing Symptoms: A Practical Way to Diagnose Hidden Causes


Leave a Reply