When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Sun, Bian, Wang, Zhenjian, de la Torre, Orvill, Wang, Zirui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
por: Barnes, Jeremy, et al.
Publicado: (2025)
por: Barnes, Jeremy, et al.
Publicado: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
por: Wang, Cheng, et al.
Publicado: (2025)
por: Wang, Cheng, et al.
Publicado: (2025)
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
por: Zhong, Qishuai, et al.
Publicado: (2025)
por: Zhong, Qishuai, et al.
Publicado: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
por: Yang, Langqi, et al.
Publicado: (2025)
por: Yang, Langqi, et al.
Publicado: (2025)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
por: Kumar, Sandeep, et al.
Publicado: (2026)
por: Kumar, Sandeep, et al.
Publicado: (2026)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
por: Sun, Kun, et al.
Publicado: (2024)
por: Sun, Kun, et al.
Publicado: (2024)
Metric-Fair Prompting: Treating Similar Samples Similarly
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
por: Zhou, Xin, et al.
Publicado: (2025)
por: Zhou, Xin, et al.
Publicado: (2025)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
por: Saha, Swarnadeep, et al.
Publicado: (2025)
por: Saha, Swarnadeep, et al.
Publicado: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
por: Wang, Yidong, et al.
Publicado: (2025)
por: Wang, Yidong, et al.
Publicado: (2025)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
por: Cao, Lang
Publicado: (2023)
por: Cao, Lang
Publicado: (2023)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025)
por: Collot, Stephane, et al.
Publicado: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
por: Ryan, Michael J., et al.
Publicado: (2025)
por: Ryan, Michael J., et al.
Publicado: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
por: Bologna, Federica, et al.
Publicado: (2026)
por: Bologna, Federica, et al.
Publicado: (2026)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
por: Wang, Zilong, et al.
Publicado: (2024)
por: Wang, Zilong, et al.
Publicado: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
por: Wang, Yidong, et al.
Publicado: (2023)
por: Wang, Yidong, et al.
Publicado: (2023)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
por: Shi, Zhichao, et al.
Publicado: (2025)
por: Shi, Zhichao, et al.
Publicado: (2025)
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
por: Yao, Zonghai, et al.
Publicado: (2024)
por: Yao, Zonghai, et al.
Publicado: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
por: Cho, Yousang, et al.
Publicado: (2025)
por: Cho, Yousang, et al.
Publicado: (2025)
A Survey on LLM-as-a-Judge
por: Gu, Jiawei, et al.
Publicado: (2024)
por: Gu, Jiawei, et al.
Publicado: (2024)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
por: Zhao, Zirui, et al.
Publicado: (2024)
por: Zhao, Zirui, et al.
Publicado: (2024)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
por: Najera, Aisha, et al.
Publicado: (2026)
por: Najera, Aisha, et al.
Publicado: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
por: Thakur, Aman Singh, et al.
Publicado: (2024)
por: Thakur, Aman Singh, et al.
Publicado: (2024)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
por: Das, Trisha, et al.
Publicado: (2024)
por: Das, Trisha, et al.
Publicado: (2024)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
por: GUO, Jiaxin, et al.
Publicado: (2025)
por: GUO, Jiaxin, et al.
Publicado: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
por: Ramprasad, Sanjana, et al.
Publicado: (2024)
por: Ramprasad, Sanjana, et al.
Publicado: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025)
por: Tong, Terry, et al.
Publicado: (2025)
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
por: Xu, Shaochen, et al.
Publicado: (2024)
por: Xu, Shaochen, et al.
Publicado: (2024)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
por: Ghanem, Hussam, et al.
Publicado: (2025)
por: Ghanem, Hussam, et al.
Publicado: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
por: Shi, Lin, et al.
Publicado: (2024)
por: Shi, Lin, et al.
Publicado: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
por: Enguehard, Joseph, et al.
Publicado: (2025)
por: Enguehard, Joseph, et al.
Publicado: (2025)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
por: Reese, May Lynn, et al.
Publicado: (2026)
por: Reese, May Lynn, et al.
Publicado: (2026)
MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes
por: Wang, Yuheng, et al.
Publicado: (2025)
por: Wang, Yuheng, et al.
Publicado: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
por: Feng, Yuanning, et al.
Publicado: (2025)
por: Feng, Yuanning, et al.
Publicado: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
por: Lai, Peng, et al.
Publicado: (2026)
por: Lai, Peng, et al.
Publicado: (2026)
Ejemplares similares
-
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025) -
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
por: Barnes, Jeremy, et al.
Publicado: (2025) -
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
por: Wang, Cheng, et al.
Publicado: (2025) -
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
por: Zhong, Qishuai, et al.
Publicado: (2025) -
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
por: Yang, Langqi, et al.
Publicado: (2025)