Dynamic Traceback Learning for Medical Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Shuchang, Meng, Mingyuan, Li, Mingjian, Feng, Dagan, Naseem, Usman, Kim, Jinman
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917049543950336
author Ye, Shuchang
Meng, Mingyuan
Li, Mingjian
Feng, Dagan
Naseem, Usman
Kim, Jinman
author_facet Ye, Shuchang
Meng, Mingyuan
Li, Mingjian
Feng, Dagan
Naseem, Usman
Kim, Jinman
contents Automated medical report generation has demonstrated the potential to significantly reduce the workload associated with time-consuming medical reporting. Recent generative representation learning methods have shown promise in integrating vision and language modalities for medical report generation. However, when trained end-to-end and applied directly to medical image-to-text generation, they face two significant challenges: i) difficulty in accurately capturing subtle yet crucial pathological details, and ii) reliance on both visual and textual inputs during inference, leading to performance degradation in zero-shot inference when only images are available. To address these challenges, this study proposes a novel multimodal dynamic traceback learning framework (DTrace). Specifically, we introduce a traceback mechanism to supervise the semantic validity of generated content and a dynamic learning strategy to adapt to various proportions of image and text input, enabling text generation without strong reliance on the input from both modalities during inference. The learning of cross-modal knowledge is enhanced by supervising the model to recover masked semantic information from a complementary counterpart. Extensive experiments conducted on two benchmark datasets, IU-Xray and MIMIC-CXR, demonstrate that the proposed DTrace framework outperforms state-of-the-art methods for medical report generation.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13267
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Traceback Learning for Medical Report Generation
Ye, Shuchang
Meng, Mingyuan
Li, Mingjian
Feng, Dagan
Naseem, Usman
Kim, Jinman
Computer Vision and Pattern Recognition
Automated medical report generation has demonstrated the potential to significantly reduce the workload associated with time-consuming medical reporting. Recent generative representation learning methods have shown promise in integrating vision and language modalities for medical report generation. However, when trained end-to-end and applied directly to medical image-to-text generation, they face two significant challenges: i) difficulty in accurately capturing subtle yet crucial pathological details, and ii) reliance on both visual and textual inputs during inference, leading to performance degradation in zero-shot inference when only images are available. To address these challenges, this study proposes a novel multimodal dynamic traceback learning framework (DTrace). Specifically, we introduce a traceback mechanism to supervise the semantic validity of generated content and a dynamic learning strategy to adapt to various proportions of image and text input, enabling text generation without strong reliance on the input from both modalities during inference. The learning of cross-modal knowledge is enhanced by supervising the model to recover masked semantic information from a complementary counterpart. Extensive experiments conducted on two benchmark datasets, IU-Xray and MIMIC-CXR, demonstrate that the proposed DTrace framework outperforms state-of-the-art methods for medical report generation.
title Dynamic Traceback Learning for Medical Report Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.13267