How AI Explanations Undermine Human Decision-Making
- Artificial intelligence tools designed to screen innovation proposals can inadvertently persuade human evaluators to reject promising projects, according to a study published by researchers associated with Harvard Business...
- In the study, researchers tested the influence of AI recommender tools on 228 experienced evaluators who assessed nearly 50 submissions to an MIT challenge.
- Our findings reveal that LLM explanations do not necessarily improve decision-making.
Artificial intelligence tools designed to screen innovation proposals can inadvertently persuade human evaluators to reject promising projects, according to a study published by researchers associated with Harvard Business School, MIT, and the University of Washington. The experiment revealed that when large language models provide written explanations for their pass-fail recommendations, humans are more likely to defer to incorrect AI decisions rather than relying on independent judgment.
How AI Rationales Suppress Independent Human Judgment
In the study, researchers tested the influence of AI recommender tools on 228 experienced evaluators who assessed nearly 50 submissions to an MIT challenge. The evaluators reviewed submissions under three distinct conditions: human-only proposals without AI assistance, large language model evaluations featuring written rationales, and black-box AI recommendations that offered a simple pass-fail decision with no accompanying explanation. Their choices were subsequently measured against a baseline established by four human experts.
Overall, evaluators accepted large language model recommendations 67 percent of the time. While they agreed with both black-box and narrative AI decisions roughly 75 percent of the time, they only matched human expert decisions 54 percent of the time. Seemingly counterintuitively, black-box recommendations improved the overall quality of decisions by aligning them more closely with human experts, whereas recommendations accompanied by narratives failed to do so. When given an automated recommendation to reject a submission along with a reason why, evaluators disproportionately concurred, which reduced false positives but substantially increased false negatives by passing on viable innovations.
Our findings reveal that LLM explanations do not necessarily improve decision-making. Effective human-AI collaboration requires designs that preserve rather than supplant independent human judgment.
The Cognitive Trap of Negativity Bias
The researchers determined that narrative explanations suppress productive overrides because large language models supply ready-made justifications that feel easy to accept, discouraging independent verification. This dynamic exploits a known human tendency called negativity bias, where people weigh negative information more heavily than positive information. Rejection functions as an active, eliminative decision that feels more accountable and consequential while maintaining the status quo, avoiding risk, and requiring no resource commitment.
Because large language models generate text that is linguistically fluent and expert-like, they create an illusion of explanatory depth. Evaluators rely on surface cues such as coherence and fluency, leading individuals to overestimate their actual understanding of a decision despite limited insight into the underlying reasoning of the model.
Implications for Enterprise AI Design
The findings suggest that organizations must exercise caution when integrating automated explanations into high-stakes decision-making workflows. While AI narratives might support conservative human decision-making in domains like compliance screening, quality control, or fraud detection, they can degrade performance in tasks like early-stage innovation screening where independent verification is critical.
Future enterprise systems could experiment with alternative designs, such as incorporating contrasting narratives that offer reasons both to accept and reject an idea, or utilizing uncertainty disclosures based on fixed thresholds rather than binary pass-fail judgments. Ultimately, the researchers advised that organizations should treat artificial intelligence explanations as behavioral interventions whose effects depend entirely on how evaluators process information under uncertainty.
