Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Yousang, Choi, Key-Sun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
von: Kim, Hyunjong, et al.
Veröffentlicht: (2025)
von: Kim, Hyunjong, et al.
Veröffentlicht: (2025)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
von: Balamurali, Sai Shridhar, et al.
Veröffentlicht: (2025)
von: Balamurali, Sai Shridhar, et al.
Veröffentlicht: (2025)
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
von: Daynauth, Roland, et al.
Veröffentlicht: (2025)
von: Daynauth, Roland, et al.
Veröffentlicht: (2025)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
von: Kang, Yipeng, et al.
Veröffentlicht: (2024)
von: Kang, Yipeng, et al.
Veröffentlicht: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
von: Li, Zongjie, et al.
Veröffentlicht: (2023)
von: Li, Zongjie, et al.
Veröffentlicht: (2023)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
von: Sun, Bian, et al.
Veröffentlicht: (2026)
von: Sun, Bian, et al.
Veröffentlicht: (2026)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
von: Toker, Gilat, et al.
Veröffentlicht: (2026)
von: Toker, Gilat, et al.
Veröffentlicht: (2026)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
von: Kim, Juyeon, et al.
Veröffentlicht: (2024)
von: Kim, Juyeon, et al.
Veröffentlicht: (2024)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
von: Lee, Woongkyu, et al.
Veröffentlicht: (2025)
von: Lee, Woongkyu, et al.
Veröffentlicht: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis Evaluation
von: Yang, Soyoung, et al.
Veröffentlicht: (2024)
von: Yang, Soyoung, et al.
Veröffentlicht: (2024)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
von: Gao, Joshua, et al.
Veröffentlicht: (2025)
von: Gao, Joshua, et al.
Veröffentlicht: (2025)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
PLEX: Perturbation-free Local Explanations for LLM-Based Text Classification
von: Rahulamathavan, Yogachandran, et al.
Veröffentlicht: (2025)
von: Rahulamathavan, Yogachandran, et al.
Veröffentlicht: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Explanation, Debate, Align: A Weak-to-Strong Framework for Language Model Generalization
von: Zakershahrak, Mehrdad, et al.
Veröffentlicht: (2024)
von: Zakershahrak, Mehrdad, et al.
Veröffentlicht: (2024)
Support-Contra Asymmetry in LLM Explanations
von: Patil, Avinash
Veröffentlicht: (2025)
von: Patil, Avinash
Veröffentlicht: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
von: Yang, Gao, et al.
Veröffentlicht: (2025)
von: Yang, Gao, et al.
Veröffentlicht: (2025)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
von: Li, Yingshu, et al.
Veröffentlicht: (2025)
von: Li, Yingshu, et al.
Veröffentlicht: (2025)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
von: Yun, Hye Sun, et al.
Veröffentlicht: (2026)
von: Yun, Hye Sun, et al.
Veröffentlicht: (2026)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments
von: Domnich, Marharyta, et al.
Veröffentlicht: (2024)
von: Domnich, Marharyta, et al.
Veröffentlicht: (2024)
Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
von: Shen, Xiangqing, et al.
Veröffentlicht: (2025)
von: Shen, Xiangqing, et al.
Veröffentlicht: (2025)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2025)
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2025)
LLM Self-Explanations Fail Semantic Invariance
von: Szeider, Stefan
Veröffentlicht: (2026)
von: Szeider, Stefan
Veröffentlicht: (2026)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
STRUX: An LLM for Decision-Making with Structured Explanations
von: Lu, Yiming, et al.
Veröffentlicht: (2024)
von: Lu, Yiming, et al.
Veröffentlicht: (2024)
Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen
von: Zainaldin, James L., et al.
Veröffentlicht: (2026)
von: Zainaldin, James L., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
von: Kim, Hyunjong, et al.
Veröffentlicht: (2025) -
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
von: Sun, Peng, et al.
Veröffentlicht: (2026) -
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
von: Balamurali, Sai Shridhar, et al.
Veröffentlicht: (2025) -
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
von: Tang, Hao, et al.
Veröffentlicht: (2024) -
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)