Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kenaston, Matthew W., Ayub, Umair, Parmar, Mihir, Anjum, Muhammad Umair, Naqvi, Syed Arsalan Ahmed, Kumar, Priya, Rawal, Samarth, Chaudhuri, Aadel A., Zakharia, Yousef, Heath, Elizabeth I., Bekaii-Saab, Tanios S., Tao, Cui, Van Allen, Eliezer M., Zhou, Ben, Choi, YooJung, Baral, Chitta, Riaz, Irbaz Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911287202545664
author Kenaston, Matthew W.
Ayub, Umair
Parmar, Mihir
Anjum, Muhammad Umair
Naqvi, Syed Arsalan Ahmed
Kumar, Priya
Rawal, Samarth
Chaudhuri, Aadel A.
Zakharia, Yousef
Heath, Elizabeth I.
Bekaii-Saab, Tanios S.
Tao, Cui
Van Allen, Eliezer M.
Zhou, Ben
Choi, YooJung
Baral, Chitta
Riaz, Irbaz Bin
author_facet Kenaston, Matthew W.
Ayub, Umair
Parmar, Mihir
Anjum, Muhammad Umair
Naqvi, Syed Arsalan Ahmed
Kumar, Priya
Rawal, Samarth
Chaudhuri, Aadel A.
Zakharia, Yousef
Heath, Elizabeth I.
Bekaii-Saab, Tanios S.
Tao, Cui
Van Allen, Eliezer M.
Zhou, Ben
Choi, YooJung
Baral, Chitta
Riaz, Irbaz Bin
contents Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
Kenaston, Matthew W.
Ayub, Umair
Parmar, Mihir
Anjum, Muhammad Umair
Naqvi, Syed Arsalan Ahmed
Kumar, Priya
Rawal, Samarth
Chaudhuri, Aadel A.
Zakharia, Yousef
Heath, Elizabeth I.
Bekaii-Saab, Tanios S.
Tao, Cui
Van Allen, Eliezer M.
Zhou, Ben
Choi, YooJung
Baral, Chitta
Riaz, Irbaz Bin
Computation and Language
Artificial Intelligence
cs.CL
I.2.7; I.2.1; J.3
Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.
title Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
topic Computation and Language
Artificial Intelligence
cs.CL
I.2.7; I.2.1; J.3
url https://arxiv.org/abs/2511.20680