Superhuman performance of a large language model on the reasoning tasks of a physician
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913871566995456 |
|---|---|
| author | Brodeur, Peter G. Buckley, Thomas A. Kanjee, Zahir Goh, Ethan Ling, Evelyn Bin Jain, Priyank Cabral, Stephanie Abdulnour, Raja-Elie Haimovich, Adrian D. Freed, Jason A. Olson, Andrew Morgan, Daniel J. Hom, Jason Gallo, Robert McCoy, Liam G. Mombini, Haadi Lucas, Christopher Fotoohi, Misha Gwiazdon, Matthew Restifo, Daniele Restrepo, Daniel Horvitz, Eric Chen, Jonathan Manrai, Arjun K. Rodman, Adam |
| author_facet | Brodeur, Peter G. Buckley, Thomas A. Kanjee, Zahir Goh, Ethan Ling, Evelyn Bin Jain, Priyank Cabral, Stephanie Abdulnour, Raja-Elie Haimovich, Adrian D. Freed, Jason A. Olson, Andrew Morgan, Daniel J. Hom, Jason Gallo, Robert McCoy, Liam G. Mombini, Haadi Lucas, Christopher Fotoohi, Misha Gwiazdon, Matthew Restifo, Daniele Restrepo, Daniel Horvitz, Eric Chen, Jonathan Manrai, Arjun K. Rodman, Adam |
| contents | A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_10849 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Superhuman performance of a large language model on the reasoning tasks of a physician Brodeur, Peter G. Buckley, Thomas A. Kanjee, Zahir Goh, Ethan Ling, Evelyn Bin Jain, Priyank Cabral, Stephanie Abdulnour, Raja-Elie Haimovich, Adrian D. Freed, Jason A. Olson, Andrew Morgan, Daniel J. Hom, Jason Gallo, Robert McCoy, Liam G. Mombini, Haadi Lucas, Christopher Fotoohi, Misha Gwiazdon, Matthew Restifo, Daniele Restrepo, Daniel Horvitz, Eric Chen, Jonathan Manrai, Arjun K. Rodman, Adam Artificial Intelligence Computation and Language A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials. |
| title | Superhuman performance of a large language model on the reasoning tasks of a physician |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2412.10849 |