Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
Fuente:
arXiv
Saved in:
| Main Authors: | DeLucia, Alexandra, Huang, Heyuan, Joshi, Sonal, Yarmohammadi, Mahsa, Hassoon, Ahmed, Dredze, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025)
by: Huang, Heyuan, et al.
Published: (2025)
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
by: DeLucia, Alexandra, et al.
Published: (2025)
by: DeLucia, Alexandra, et al.
Published: (2025)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
by: Jahara, Fatima, et al.
Published: (2025)
by: Jahara, Fatima, et al.
Published: (2025)
Transferring Fairness using Multi-Task Learning with Limited Demographic Information
by: Aguirre, Carlos, et al.
Published: (2023)
by: Aguirre, Carlos, et al.
Published: (2023)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
by: Ehara, Yo
Published: (2026)
by: Ehara, Yo
Published: (2026)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
by: Liu, Jonathan, et al.
Published: (2025)
by: Liu, Jonathan, et al.
Published: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025)
by: Han, Steve, et al.
Published: (2025)
Anti-LM Decoding for Zero-shot In-context Machine Translation
by: Sia, Suzanna, et al.
Published: (2023)
by: Sia, Suzanna, et al.
Published: (2023)
Same Patients, Different Health Care Systems—Revisited. Geriatric Care Models in the U.S., Canada, and Europe
by: Nathalie van der Velde, et al.
Published: (2026)
by: Nathalie van der Velde, et al.
Published: (2026)
The Value of Disagreement in AI Design, Evaluation, and Alignment
by: Fazelpour, Sina, et al.
Published: (2025)
by: Fazelpour, Sina, et al.
Published: (2025)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Relational Mediators: LLM Chatbots as Boundary Objects in Psychotherapy
by: Quan, Jiatao, et al.
Published: (2025)
by: Quan, Jiatao, et al.
Published: (2025)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Review Helpfulness Scores vs. Review Unhelpfulness Scores: Two Sides of the Same Coin or Different Coins?
by: Yu, Yinan, et al.
Published: (2024)
by: Yu, Yinan, et al.
Published: (2024)
SARHAchat: An LLM-Based Chatbot for Sexual and Reproductive Health Counseling
by: Yang, Jiaye, et al.
Published: (2025)
by: Yang, Jiaye, et al.
Published: (2025)
Verdict: A Library for Scaling Judge-Time Compute
by: Kalra, Nimit, et al.
Published: (2025)
by: Kalra, Nimit, et al.
Published: (2025)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Understanding Privacy Norms Around LLM-Based Chatbots: A Contextual Integrity Perspective
by: Tran, Sarah, et al.
Published: (2025)
by: Tran, Sarah, et al.
Published: (2025)
A Mixed-Methods Evaluation of LLM-Based Chatbots for Menopause
by: Deva, Roshini, et al.
Published: (2025)
by: Deva, Roshini, et al.
Published: (2025)
LLM Chatbots in High School Programming: Exploring Behaviors and Interventions
by: Torre, Manuel Valle, et al.
Published: (2025)
by: Torre, Manuel Valle, et al.
Published: (2025)
A Cleaner Production Scheduling Model with Green Investment and Pandemic Effects
by: S. Priyan, et al.
Published: (2024)
by: S. Priyan, et al.
Published: (2024)
Echoes of Disagreement: Measuring Disparity in Social Consensus
by: Papachristou, Marios, et al.
Published: (2025)
by: Papachristou, Marios, et al.
Published: (2025)
EXAGREE: Mitigating Explanation Disagreement with Stakeholder-Aligned Models
by: Li, Sichao, et al.
Published: (2024)
by: Li, Sichao, et al.
Published: (2024)
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
We're Different, We're the Same: Creative Homogeneity Across LLMs
by: Wenger, Emily, et al.
Published: (2025)
by: Wenger, Emily, et al.
Published: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
by: He, Peng, et al.
Published: (2026)
by: He, Peng, et al.
Published: (2026)
WaLLM -- Insights from an LLM-Powered Chatbot deployment via WhatsApp
by: Eltigani, Hiba, et al.
Published: (2025)
by: Eltigani, Hiba, et al.
Published: (2025)
Mitigating the Carbon Footprint of Chatbots as Consumers
by: Ruf, Boris, et al.
Published: (2025)
by: Ruf, Boris, et al.
Published: (2025)
Behavior and Sublinear Algorithm for Opinion Disagreement on Noisy Social Networks
by: Xu, Wanyue, et al.
Published: (2026)
by: Xu, Wanyue, et al.
Published: (2026)
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
by: Li, Lingyao, et al.
Published: (2026)
by: Li, Lingyao, et al.
Published: (2026)
VayuBuddy: an LLM-Powered Chatbot to Democratize Air Quality Insights
by: Patel, Zeel B, et al.
Published: (2024)
by: Patel, Zeel B, et al.
Published: (2024)
Reply to: Comment on: The Halo Effect: Perceptions of Information Privacy Among Healthcare Chatbot Users
by: Matthew DeCamp, et al.
Published: (2025)
by: Matthew DeCamp, et al.
Published: (2025)
Trust as a Situated User State in Social LLM-Based Chatbots: A Longitudinal Study of Snapchat's My AI
by: Landerberg, Annie, et al.
Published: (2026)
by: Landerberg, Annie, et al.
Published: (2026)
Clinician and Informant Report of Neuropsychiatric Symptoms in Dementia
by: Carolyn W. Zhu, et al.
Published: (2026)
by: Carolyn W. Zhu, et al.
Published: (2026)
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
by: Aftahee, Sabik, et al.
Published: (2025)
by: Aftahee, Sabik, et al.
Published: (2025)
JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
by: Razumenko, Itay, et al.
Published: (2026)
by: Razumenko, Itay, et al.
Published: (2026)
Learn, Explore and Reflect by Chatting: Understanding the Value of an LLM-Based Voting Advice Application Chatbot
by: Zhu, Jianlong, et al.
Published: (2025)
by: Zhu, Jianlong, et al.
Published: (2025)
Can Consumer Chatbots Reason? A Student-Led Field Experiment Embedded in an "AI-for-All" Undergraduate Course
by: Shehu, Amarda, et al.
Published: (2025)
by: Shehu, Amarda, et al.
Published: (2025)
Tailoring Chatbots for Higher Education: Some Insights and Experiences
by: Kortemeyer, Gerd
Published: (2024)
by: Kortemeyer, Gerd
Published: (2024)
Similar Items
-
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
by: Huang, Heyuan, et al.
Published: (2025) -
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
by: DeLucia, Alexandra, et al.
Published: (2025) -
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
by: Jahara, Fatima, et al.
Published: (2025) -
Transferring Fairness using Multi-Task Learning with Limited Demographic Information
by: Aguirre, Carlos, et al.
Published: (2023) -
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
by: Ehara, Yo
Published: (2026)