RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yingshu, Liu, Yunyi, Liu, Lingqiao, Wang, Lei, Zhou, Luping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MRScore: Evaluating Radiology Report Generation with LLM-based Reward System
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
KARGEN: Knowledge-enhanced Automated Radiology Report Generation Using Large Language Models
von: Li, Yingshu, et al.
Veröffentlicht: (2024)
von: Li, Yingshu, et al.
Veröffentlicht: (2024)
A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
von: Zhou, Shaoyang, et al.
Veröffentlicht: (2025)
von: Zhou, Shaoyang, et al.
Veröffentlicht: (2025)
S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
von: Li, Yingshu, et al.
Veröffentlicht: (2025)
von: Li, Yingshu, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
von: Li, Yingshu, et al.
Veröffentlicht: (2023)
von: Li, Yingshu, et al.
Veröffentlicht: (2023)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling
von: Lin, Ying-Jia, et al.
Veröffentlicht: (2026)
von: Lin, Ying-Jia, et al.
Veröffentlicht: (2026)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
von: Shetty, Saisha Pradeep, et al.
Veröffentlicht: (2026)
von: Shetty, Saisha Pradeep, et al.
Veröffentlicht: (2026)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
von: Wang, Haofeng, et al.
Veröffentlicht: (2025)
von: Wang, Haofeng, et al.
Veröffentlicht: (2025)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
von: He, Zhonghao, et al.
Veröffentlicht: (2025)
von: He, Zhonghao, et al.
Veröffentlicht: (2025)
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology
von: Fytas, Panagiotis, et al.
Veröffentlicht: (2024)
von: Fytas, Panagiotis, et al.
Veröffentlicht: (2024)
RadEx: A Framework for Structured Information Extraction from Radiology Reports based on Large Language Models
von: Reichenpfader, Daniel, et al.
Veröffentlicht: (2024)
von: Reichenpfader, Daniel, et al.
Veröffentlicht: (2024)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
von: Sun, Chengbo, et al.
Veröffentlicht: (2025)
von: Sun, Chengbo, et al.
Veröffentlicht: (2025)
Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation
von: Wang, Xinyi, et al.
Veröffentlicht: (2026)
von: Wang, Xinyi, et al.
Veröffentlicht: (2026)
RadFabric: Agentic AI System with Reasoning Capability for Radiology
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses
von: Ren, Xuan, et al.
Veröffentlicht: (2024)
von: Ren, Xuan, et al.
Veröffentlicht: (2024)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
GREEN: Generative Radiology Report Evaluation and Error Notation
von: Ostmeier, Sophie, et al.
Veröffentlicht: (2024)
von: Ostmeier, Sophie, et al.
Veröffentlicht: (2024)
MGH Radiology Llama: A Llama 3 70B Model for Radiology
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
von: Shi, Yucheng, et al.
Veröffentlicht: (2024)
Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
von: Xu, Shaochen, et al.
Veröffentlicht: (2024)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
Radiology-GPT: A Large Language Model for Radiology
von: Liu, Zhengliang, et al.
Veröffentlicht: (2023)
von: Liu, Zhengliang, et al.
Veröffentlicht: (2023)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
von: Fan, Ziqing, et al.
Veröffentlicht: (2025)
von: Fan, Ziqing, et al.
Veröffentlicht: (2025)
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
von: Han, Simeng, et al.
Veröffentlicht: (2024)
von: Han, Simeng, et al.
Veröffentlicht: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
von: Jiang, Yuyang, et al.
Veröffentlicht: (2025)
von: Jiang, Yuyang, et al.
Veröffentlicht: (2025)
ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation
von: Lee, Zhicheng, et al.
Veröffentlicht: (2025)
von: Lee, Zhicheng, et al.
Veröffentlicht: (2025)
Skywork Open Reasoner 1 Technical Report
von: He, Jujie, et al.
Veröffentlicht: (2025)
von: He, Jujie, et al.
Veröffentlicht: (2025)
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
von: Zhu, Qingqing, et al.
Veröffentlicht: (2024)
FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in Large Language Models
von: Li, Yiyuan, et al.
Veröffentlicht: (2024)
von: Li, Yiyuan, et al.
Veröffentlicht: (2024)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
Teach-to-Reason with Scoring: Self-Explainable Rationale-Driven Multi-Trait Essay Scoring
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MRScore: Evaluating Radiology Report Generation with LLM-based Reward System
von: Liu, Yunyi, et al.
Veröffentlicht: (2024) -
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
von: Liu, Yunyi, et al.
Veröffentlicht: (2024) -
KARGEN: Knowledge-enhanced Automated Radiology Report Generation Using Large Language Models
von: Li, Yingshu, et al.
Veröffentlicht: (2024) -
A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
von: Zhou, Shaoyang, et al.
Veröffentlicht: (2025) -
S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
von: Li, Yingshu, et al.
Veröffentlicht: (2025)