FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Alyssa, Banerjee, Oishi, Wu, Kay, Reis, Eduardo Pontes, Rajpurkar, Pranav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
by: Banerjee, Oishi, et al.
Published: (2024)
by: Banerjee, Oishi, et al.
Published: (2024)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports
by: Hardy, Romain, et al.
Published: (2024)
by: Hardy, Romain, et al.
Published: (2024)
Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation
by: Zhang, Zhenxuan, et al.
Published: (2025)
by: Zhang, Zhenxuan, et al.
Published: (2025)
ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports
by: Rao, Vishwanatha M., et al.
Published: (2024)
by: Rao, Vishwanatha M., et al.
Published: (2024)
ReXamine-Global: A Framework for Uncovering Inconsistencies in Radiology Report Generation Metrics
by: Banerjee, Oishi, et al.
Published: (2024)
by: Banerjee, Oishi, et al.
Published: (2024)
RadFlag: A Black-Box Hallucination Detection Method for Medical Vision Language Models
by: Zhang, Serena, et al.
Published: (2024)
by: Zhang, Serena, et al.
Published: (2024)
CRG Score: A Distribution-Aware Clinical Metric for Radiology Report Generation
by: Hamamci, Ibrahim Ethem, et al.
Published: (2025)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2025)
RadBARTsum: Domain Specific Adaption of Denoising Sequence-to-Sequence Models for Abstractive Radiology Report Summarization
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
by: Zhou, Sitong, et al.
Published: (2026)
by: Zhou, Sitong, et al.
Published: (2026)
RadFabric: Agentic AI System with Reasoning Capability for Radiology
by: Chen, Wenting, et al.
Published: (2025)
by: Chen, Wenting, et al.
Published: (2025)
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
by: Choi, Hongseok, et al.
Published: (2026)
by: Choi, Hongseok, et al.
Published: (2026)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
Automated Structured Radiology Report Generation
by: Delbrouck, Jean-Benoit, et al.
Published: (2025)
by: Delbrouck, Jean-Benoit, et al.
Published: (2025)
RadPhi-3: Small Language Models for Radiology
by: Ranjit, Mercy, et al.
Published: (2024)
by: Ranjit, Mercy, et al.
Published: (2024)
Ran Score: a LLM-based Evaluation Score for Radiology Report Generation
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
VideoScore2: Think before You Score in Generative Video Evaluation
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling
by: Lin, Ying-Jia, et al.
Published: (2026)
by: Lin, Ying-Jia, et al.
Published: (2026)
The Impact of AI Assistance on Radiology Reporting: A Pilot Study Using Simulated AI Draft Reports
by: Acosta, Julián N., et al.
Published: (2024)
by: Acosta, Julián N., et al.
Published: (2024)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2026)
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology
by: Fytas, Panagiotis, et al.
Published: (2024)
by: Fytas, Panagiotis, et al.
Published: (2024)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
by: Gupta, Pankaj, et al.
Published: (2026)
by: Gupta, Pankaj, et al.
Published: (2026)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
by: Shen, Xiaoxian, et al.
Published: (2026)
by: Shen, Xiaoxian, et al.
Published: (2026)
Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework
by: Wu, Jinge, et al.
Published: (2024)
by: Wu, Jinge, et al.
Published: (2024)
GREEN: Generative Radiology Report Evaluation and Error Notation
by: Ostmeier, Sophie, et al.
Published: (2024)
by: Ostmeier, Sophie, et al.
Published: (2024)
RadEx: A Framework for Structured Information Extraction from Radiology Reports based on Large Language Models
by: Reichenpfader, Daniel, et al.
Published: (2024)
by: Reichenpfader, Daniel, et al.
Published: (2024)
Autoregressive Score Generation for Multi-trait Essay Scoring
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
RaTEScore: A Metric for Radiology Report Generation
by: Zhao, Weike, et al.
Published: (2024)
by: Zhao, Weike, et al.
Published: (2024)
Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings
by: Mahmood, Razi, et al.
Published: (2024)
by: Mahmood, Razi, et al.
Published: (2024)
FRACTAL: Fine-Grained Scoring from Aggregate Text Labels
by: Makhija, Yukti, et al.
Published: (2024)
by: Makhija, Yukti, et al.
Published: (2024)
Two-Pronged Human Evaluation of ChatGPT Self-Correction in Radiology Report Simplification
by: Yang, Ziyu, et al.
Published: (2024)
by: Yang, Ziyu, et al.
Published: (2024)
X-ray Made Simple: Lay Radiology Report Generation and Robust Evaluation
by: Zhao, Kun, et al.
Published: (2024)
by: Zhao, Kun, et al.
Published: (2024)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Calibrated Confidence Expression for Radiology Report Generation
by: Bani-Harouni, David, et al.
Published: (2026)
by: Bani-Harouni, David, et al.
Published: (2026)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
by: Sun, Chengbo, et al.
Published: (2025)
by: Sun, Chengbo, et al.
Published: (2025)
MRScore: Evaluating Radiology Report Generation with LLM-based Reward System
by: Liu, Yunyi, et al.
Published: (2024)
by: Liu, Yunyi, et al.
Published: (2024)
Similar Items
-
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
by: Banerjee, Oishi, et al.
Published: (2024) -
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025) -
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
by: Zhang, Xiaoman, et al.
Published: (2024) -
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026) -
ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports
by: Hardy, Romain, et al.
Published: (2024)