Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuan, Gao, Shujian, Liu, Jiaxiang, Jiang, Songtao, Xia, Haoxiang, Zhang, Xiaotian, Kang, Zhaolu, Wang, Yemin, Liu, Zuozhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915648842498048
author Wang, Yuan
Gao, Shujian
Liu, Jiaxiang
Jiang, Songtao
Xia, Haoxiang
Zhang, Xiaotian
Kang, Zhaolu
Wang, Yemin
Liu, Zuozhu
author_facet Wang, Yuan
Gao, Shujian
Liu, Jiaxiang
Jiang, Songtao
Xia, Haoxiang
Zhang, Xiaotian
Kang, Zhaolu
Wang, Yemin
Liu, Zuozhu
contents Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 12.1% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02710
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
Wang, Yuan
Gao, Shujian
Liu, Jiaxiang
Jiang, Songtao
Xia, Haoxiang
Zhang, Xiaotian
Kang, Zhaolu
Wang, Yemin
Liu, Zuozhu
Computational Engineering, Finance, and Science
Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 12.1% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.
title Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2512.02710