Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yanjun, Myers, Skatje, Chen, Shan, Dligach, Dmitriy, Miller, Timothy A, Bitterman, Danielle, Chen, Guanhua, Mayampurath, Anoop, Churpek, Matthew, Afshar, Majid
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912109532545024
author Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Chen, Guanhua
Mayampurath, Anoop
Churpek, Matthew
Afshar, Majid
author_facet Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Chen, Guanhua
Mayampurath, Anoop
Churpek, Matthew
Afshar, Majid
contents Large language models (LLMs) are being explored for diagnostic decision support, yet their ability to estimate pre-test probabilities, vital for clinical decision-making, remains limited. This study evaluates two LLMs, Mistral-7B and Llama3-70B, using structured electronic health record data on three diagnosis tasks. We examined three current methods of extracting LLM probability estimations and revealed their limitations. We aim to highlight the need for improved techniques in LLM confidence estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04962
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
Gao, Yanjun
Myers, Skatje
Chen, Shan
Dligach, Dmitriy
Miller, Timothy A
Bitterman, Danielle
Chen, Guanhua
Mayampurath, Anoop
Churpek, Matthew
Afshar, Majid
Artificial Intelligence
Computation and Language
Large language models (LLMs) are being explored for diagnostic decision support, yet their ability to estimate pre-test probabilities, vital for clinical decision-making, remains limited. This study evaluates two LLMs, Mistral-7B and Llama3-70B, using structured electronic health record data on three diagnosis tasks. We examined three current methods of extracting LLM probability estimations and revealed their limitations. We aim to highlight the need for improved techniques in LLM confidence estimation.
title Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.04962