Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
Fuente:
arXiv
Saved in:
| Main Authors: | Testoni, Alberto, Calixto, Iacer |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
by: Testoni, Alberto, et al.
Published: (2025)
by: Testoni, Alberto, et al.
Published: (2025)
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
by: Chakravarty, Abhirup, et al.
Published: (2025)
by: Chakravarty, Abhirup, et al.
Published: (2025)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
LLM for Everyone: Representing the Underrepresented in Large Language Models
by: Cahyawijaya, Samuel
Published: (2024)
by: Cahyawijaya, Samuel
Published: (2024)
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026)
by: Yun, Hye Sun, et al.
Published: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
by: Miranda, Michele, et al.
Published: (2026)
by: Miranda, Michele, et al.
Published: (2026)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Fact-Level Confidence Calibration and Self-Correction
by: Yuan, Yige, et al.
Published: (2024)
by: Yuan, Yige, et al.
Published: (2024)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
by: Lu, Yuyin, et al.
Published: (2026)
by: Lu, Yuyin, et al.
Published: (2026)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
by: Hong, Chang, et al.
Published: (2025)
by: Hong, Chang, et al.
Published: (2025)
How Confident Is the First Token? An Uncertainty-Calibrated Prompt Optimization Framework for Large Language Model Classification and Understanding
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
by: Harshavardhan
Published: (2026)
by: Harshavardhan
Published: (2026)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
by: Bavaresco, A., et al.
Published: (2024)
by: Bavaresco, A., et al.
Published: (2024)
DeCode: Decoupling Content and Delivery for Medical QA
by: Ko, Po-Jen, et al.
Published: (2026)
by: Ko, Po-Jen, et al.
Published: (2026)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
by: Geng, Jiahui, et al.
Published: (2023)
by: Geng, Jiahui, et al.
Published: (2023)
How LLMs Distort Our Written Language
by: Abdulhai, Marwa, et al.
Published: (2026)
by: Abdulhai, Marwa, et al.
Published: (2026)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
by: Torrielli, Federico, et al.
Published: (2026)
by: Torrielli, Federico, et al.
Published: (2026)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
by: Kim, Yunsoo, et al.
Published: (2024)
by: Kim, Yunsoo, et al.
Published: (2024)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
by: Alan, Ahmet Yusuf, et al.
Published: (2024)
by: Alan, Ahmet Yusuf, et al.
Published: (2024)
The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration
by: Ghosh, Sudipta, et al.
Published: (2026)
by: Ghosh, Sudipta, et al.
Published: (2026)
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
by: Zhang, Chuang, et al.
Published: (2026)
by: Zhang, Chuang, et al.
Published: (2026)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2025)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
by: Chhikara, Prateek
Published: (2025)
by: Chhikara, Prateek
Published: (2025)
From Confidence to Collapse in LLM Factual Robustness
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
PassiveQA: A Three-Action Framework for Epistemically Calibrated Question Answering via Supervised Finetuning
by: Baidya, Madhav S
Published: (2026)
by: Baidya, Madhav S
Published: (2026)
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA
by: Roy, Nirmal, et al.
Published: (2024)
by: Roy, Nirmal, et al.
Published: (2024)
LLM4Causal: Democratized Causal Tools for Everyone via Large Language Model
by: Jiang, Haitao, et al.
Published: (2023)
by: Jiang, Haitao, et al.
Published: (2023)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
by: Huang, Liangjie, et al.
Published: (2025)
by: Huang, Liangjie, et al.
Published: (2025)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
by: Yang, Hang, et al.
Published: (2024)
by: Yang, Hang, et al.
Published: (2024)
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
by: Jung, Jimin, et al.
Published: (2026)
by: Jung, Jimin, et al.
Published: (2026)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
Confidence Estimation for LLM-Based Dialogue State Tracking
by: Sun, Yi-Jyun, et al.
Published: (2024)
by: Sun, Yi-Jyun, et al.
Published: (2024)
Similar Items
-
Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
by: Testoni, Alberto, et al.
Published: (2025) -
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
by: Zhang, Zeyu, et al.
Published: (2024) -
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
by: Testoni, Alberto, et al.
Published: (2024) -
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
by: Chakravarty, Abhirup, et al.
Published: (2025) -
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)