Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
Fuente:
arXiv
Saved in:
| Main Authors: | Naderi, Nariman, Safavi-Naini, Seyed Amir Ahmad, Savage, Thomas, Atf, Zahra, Lewis, Peter, Nadkarni, Girish, Soroush, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
The challenge of uncertainty quantification of large language models in medicine
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024)
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024)
State of Abdominal CT Datasets: A Critical Review of Bias, Clinical Relevance, and Real-world Applicability
by: Danaei, Saeide, et al.
Published: (2025)
by: Danaei, Saeide, et al.
Published: (2025)
Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2026)
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2026)
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images
by: Khalafi, Mohammad Amin, et al.
Published: (2025)
by: Khalafi, Mohammad Amin, et al.
Published: (2025)
Rule-Based Moral Principles for Explaining Uncertainty in Natural Language Generation
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Is Trust Correlated With Explainability in AI? A Meta-Analysis
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Kantian-Utilitarian XAI: Meta-Explained
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
3DLAND: 3D Lesion Abdominal Anomaly Localization Dataset
by: Advand, Mehran, et al.
Published: (2026)
by: Advand, Mehran, et al.
Published: (2026)
A Physics-Informed Machine Learning Framework for Solid Boundary Treatment in Meshfree Particle Methods
by: Mehranfar, Nariman, et al.
Published: (2025)
by: Mehranfar, Nariman, et al.
Published: (2025)
Hybrid Encryption with Certified Deletion in Preprocessing Model
by: Dey, Kunal, et al.
Published: (2026)
by: Dey, Kunal, et al.
Published: (2026)
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports
by: Dorfner, Felix J., et al.
Published: (2024)
by: Dorfner, Felix J., et al.
Published: (2024)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
by: Cardoso, Gabriel Pimenta de Freitas, et al.
Published: (2026)
When Quantization Affects Confidence of Large Language Models?
by: Proskurina, Irina, et al.
Published: (2024)
by: Proskurina, Irina, et al.
Published: (2024)
Confidence Improves Self-Consistency in LLMs
by: Taubenfeld, Amir, et al.
Published: (2025)
by: Taubenfeld, Amir, et al.
Published: (2025)
Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical Tasks
by: Ateia, Samy, et al.
Published: (2024)
by: Ateia, Samy, et al.
Published: (2024)
Large Language Models versus Classical Machine Learning: Performance in COVID-19 Mortality Prediction Using High-Dimensional Tabular Data
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2024)
by: Ghaffarzadeh-Esfahani, Mohammadreza, et al.
Published: (2024)
Robust and Reusable Fuzzy Extractors for Low-entropy Rate Randomness Sources
by: Panja, Somnath, et al.
Published: (2024)
by: Panja, Somnath, et al.
Published: (2024)
Third-Party Language Model Performance Prediction from Instruction
by: Nadkarni, Rahul, et al.
Published: (2024)
by: Nadkarni, Rahul, et al.
Published: (2024)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
OpenUS: A Fully Open-Source Foundation Model for Ultrasound Image Analysis via Self-Adaptive Masked Contrastive Learning
by: Zheng, Xiaoyu, et al.
Published: (2025)
by: Zheng, Xiaoyu, et al.
Published: (2025)
An Open Source Data Contamination Report for Large Language Models
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence
by: Kazemzadeh, Houman, et al.
Published: (2026)
by: Kazemzadeh, Houman, et al.
Published: (2026)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
An Open-Source Fast Parallel Routing Approach for Commercial FPGAs
by: Zang, Xinshi, et al.
Published: (2024)
by: Zang, Xinshi, et al.
Published: (2024)
SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2026)
by: Oraki, Soroush, et al.
Published: (2026)
Serial Position Effects of Large Language Models
by: Guo, Xiaobo, et al.
Published: (2024)
by: Guo, Xiaobo, et al.
Published: (2024)
Self-calibration for Language Model Quantization and Pruning
by: Williams, Miles, et al.
Published: (2024)
by: Williams, Miles, et al.
Published: (2024)
Secure Composition of Quantum Key Distribution and Symmetric Key Encryption
by: Dey, Kunal, et al.
Published: (2025)
by: Dey, Kunal, et al.
Published: (2025)
Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching
by: Savage, Thomas
Published: (2025)
by: Savage, Thomas
Published: (2025)
PsycholexTherapy: Simulating Reasoning in Psychotherapy with Small Language Models in Persian
by: Abbasi, Mohammad Amin, et al.
Published: (2025)
by: Abbasi, Mohammad Amin, et al.
Published: (2025)
CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
by: Aghazadeh, Ehsan, et al.
Published: (2025)
by: Aghazadeh, Ehsan, et al.
Published: (2025)
Pixel-wise RL on Diffusion Models: Reinforcement Learning from Rich Feedback
by: Kordzanganeh, Mo, et al.
Published: (2024)
by: Kordzanganeh, Mo, et al.
Published: (2024)
Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
by: Koma, Arian Komaei, et al.
Published: (2026)
by: Koma, Arian Komaei, et al.
Published: (2026)
Evaluating Prompt Engineering Techniques for RAG in Small Language Models: A Multi-Hop QA Approach
by: Mohammadi, Amir Hossein, et al.
Published: (2026)
by: Mohammadi, Amir Hossein, et al.
Published: (2026)
Allegro: Open the Black Box of Commercial-Level Video Generation Model
by: Zhou, Yuan, et al.
Published: (2024)
by: Zhou, Yuan, et al.
Published: (2024)
A Feature-Level Ensemble Model for COVID-19 Identification in CXR Images using Choquet Integral and Differential Evolution Optimization
by: Takhsha, Amir Reza, et al.
Published: (2025)
by: Takhsha, Amir Reza, et al.
Published: (2025)
Self-Training Large Language Models with Confident Reasoning
by: Jang, Hyosoon, et al.
Published: (2025)
by: Jang, Hyosoon, et al.
Published: (2025)
Similar Items
-
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025) -
The challenge of uncertainty quantification of large language models in medicine
by: Atf, Zahra, et al.
Published: (2025) -
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024) -
State of Abdominal CT Datasets: A Critical Review of Bias, Clinical Relevance, and Real-world Applicability
by: Danaei, Saeide, et al.
Published: (2025) -
Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2026)