Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naderi, Nariman, Safavi-Naini, Seyed Amir Ahmad, Savage, Thomas, Atf, Zahra, Lewis, Peter, Nadkarni, Girish, Soroush, Ali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910114045231104
author Naderi, Nariman
Safavi-Naini, Seyed Amir Ahmad
Savage, Thomas
Atf, Zahra
Lewis, Peter
Nadkarni, Girish
Soroush, Ali
author_facet Naderi, Nariman
Safavi-Naini, Seyed Amir Ahmad
Savage, Thomas
Atf, Zahra
Lewis, Peter
Nadkarni, Girish
Soroush, Ali
contents This study evaluated self-reported response certainty across several large language models (GPT, Claude, Llama, Phi, Mistral, Gemini, Gemma, and Qwen) using 300 gastroenterology board-style questions. The highest-performing models (GPT-o1 preview, GPT-4o, and Claude-3.5-Sonnet) achieved Brier scores of 0.15-0.2 and AUROC of 0.6. Although newer models demonstrated improved performance, all exhibited a consistent tendency towards overconfidence. Uncertainty estimation presents a significant challenge to the safe use of LLMs in healthcare. Keywords: Large Language Models; Confidence Elicitation; Artificial Intelligence; Gastroenterology; Uncertainty Quantification
format Preprint
id arxiv_https___arxiv_org_abs_2503_18562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
Naderi, Nariman
Safavi-Naini, Seyed Amir Ahmad
Savage, Thomas
Atf, Zahra
Lewis, Peter
Nadkarni, Girish
Soroush, Ali
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
This study evaluated self-reported response certainty across several large language models (GPT, Claude, Llama, Phi, Mistral, Gemini, Gemma, and Qwen) using 300 gastroenterology board-style questions. The highest-performing models (GPT-o1 preview, GPT-4o, and Claude-3.5-Sonnet) achieved Brier scores of 0.15-0.2 and AUROC of 0.6. Although newer models demonstrated improved performance, all exhibited a consistent tendency towards overconfidence. Uncertainty estimation presents a significant challenge to the safe use of LLMs in healthcare. Keywords: Large Language Models; Confidence Elicitation; Artificial Intelligence; Gastroenterology; Uncertainty Quantification
title Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2503.18562