When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Bian, Wang, Zhenjian, de la Torre, Orvill, Wang, Zirui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
von: Zhong, Qishuai, et al.
Veröffentlicht: (2025)
von: Zhong, Qishuai, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
von: Sun, Kun, et al.
Veröffentlicht: (2024)
von: Sun, Kun, et al.
Veröffentlicht: (2024)
Metric-Fair Prompting: Treating Similar Samples Similarly
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
von: Yao, Zonghai, et al.
Veröffentlicht: (2024)
von: Yao, Zonghai, et al.
Veröffentlicht: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
von: Cho, Yousang, et al.
Veröffentlicht: (2025)
von: Cho, Yousang, et al.
Veröffentlicht: (2025)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
von: Zhao, Zirui, et al.
Veröffentlicht: (2024)
von: Zhao, Zirui, et al.
Veröffentlicht: (2024)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
von: Najera, Aisha, et al.
Veröffentlicht: (2026)
von: Najera, Aisha, et al.
Veröffentlicht: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
von: Das, Trisha, et al.
Veröffentlicht: (2024)
von: Das, Trisha, et al.
Veröffentlicht: (2024)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
von: Ghanem, Hussam, et al.
Veröffentlicht: (2025)
von: Ghanem, Hussam, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
von: Enguehard, Joseph, et al.
Veröffentlicht: (2025)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes
von: Wang, Yuheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuheng, et al.
Veröffentlicht: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
von: Lai, Peng, et al.
Veröffentlicht: (2026)
von: Lai, Peng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025) -
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025) -
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
von: Wang, Cheng, et al.
Veröffentlicht: (2025) -
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
von: Zhong, Qishuai, et al.
Veröffentlicht: (2025) -
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)