MRScore: Evaluating Radiology Report Generation with LLM-based Reward System
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yunyi, Wang, Zhanyu, Li, Yingshu, Liang, Xinyu, Liu, Lingqiao, Wang, Lei, Zhou, Luping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
by: Liu, Yunyi, et al.
Published: (2024)
by: Liu, Yunyi, et al.
Published: (2024)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
KARGEN: Knowledge-enhanced Automated Radiology Report Generation Using Large Language Models
by: Li, Yingshu, et al.
Published: (2024)
by: Li, Yingshu, et al.
Published: (2024)
S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
by: Li, Yingshu, et al.
Published: (2023)
by: Li, Yingshu, et al.
Published: (2023)
A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
by: Zhou, Shaoyang, et al.
Published: (2025)
by: Zhou, Shaoyang, et al.
Published: (2025)
I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses
by: Ren, Xuan, et al.
Published: (2024)
by: Ren, Xuan, et al.
Published: (2024)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
by: Sun, Chengbo, et al.
Published: (2025)
by: Sun, Chengbo, et al.
Published: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
GREEN: Generative Radiology Report Evaluation and Error Notation
by: Ostmeier, Sophie, et al.
Published: (2024)
by: Ostmeier, Sophie, et al.
Published: (2024)
Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports
by: Thomas, Alois, et al.
Published: (2025)
by: Thomas, Alois, et al.
Published: (2025)
Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation
by: Wang, Xinyi, et al.
Published: (2026)
by: Wang, Xinyi, et al.
Published: (2026)
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
by: Li, Zhaoyan, et al.
Published: (2026)
by: Li, Zhaoyan, et al.
Published: (2026)
Enhancing LLMs for Impression Generation in Radiology Reports through a Multi-Agent System
by: Zeng, Fang, et al.
Published: (2024)
by: Zeng, Fang, et al.
Published: (2024)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
by: Dua, Radhika, et al.
Published: (2025)
by: Dua, Radhika, et al.
Published: (2025)
You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
by: Ren, Xuan, et al.
Published: (2023)
by: Ren, Xuan, et al.
Published: (2023)
Generative Large Language Models Trained for Detecting Errors in Radiology Reports
by: Sun, Cong, et al.
Published: (2025)
by: Sun, Cong, et al.
Published: (2025)
StoryAlign: Evaluating and Training Reward Models for Story Generation
by: Xia, Haotian, et al.
Published: (2026)
by: Xia, Haotian, et al.
Published: (2026)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
MGH Radiology Llama: A Llama 3 70B Model for Radiology
by: Shi, Yucheng, et al.
Published: (2024)
by: Shi, Yucheng, et al.
Published: (2024)
Is Your LLM Outdated? A Deep Look at Temporal Generalization
by: Zhu, Chenghao, et al.
Published: (2024)
by: Zhu, Chenghao, et al.
Published: (2024)
Radiology-GPT: A Large Language Model for Radiology
by: Liu, Zhengliang, et al.
Published: (2023)
by: Liu, Zhengliang, et al.
Published: (2023)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
by: Ding, Liang
Published: (2026)
by: Ding, Liang
Published: (2026)
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
by: Jin, Song, et al.
Published: (2025)
by: Jin, Song, et al.
Published: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
by: Li, Zongjie, et al.
Published: (2023)
by: Li, Zongjie, et al.
Published: (2023)
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
by: Zhu, Qingqing, et al.
Published: (2024)
by: Zhu, Qingqing, et al.
Published: (2024)
Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports
by: Niu, Chuang, et al.
Published: (2024)
by: Niu, Chuang, et al.
Published: (2024)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
by: Peng, Hao, et al.
Published: (2025)
by: Peng, Hao, et al.
Published: (2025)
S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation
by: Gao, Jiechao, et al.
Published: (2025)
by: Gao, Jiechao, et al.
Published: (2025)
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
by: Feng, Zhaopeng, et al.
Published: (2025)
by: Feng, Zhaopeng, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026)
by: Liu, Runheng, et al.
Published: (2026)
Ran Score: a LLM-based Evaluation Score for Radiology Report Generation
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
by: Wu, Wanxing, et al.
Published: (2026)
by: Wu, Wanxing, et al.
Published: (2026)
Similar Items
-
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
by: Liu, Yunyi, et al.
Published: (2024) -
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025) -
KARGEN: Knowledge-enhanced Automated Radiology Report Generation Using Large Language Models
by: Li, Yingshu, et al.
Published: (2024) -
S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
by: Li, Yingshu, et al.
Published: (2025) -
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
by: Li, Yingshu, et al.
Published: (2023)