MedRAT: Unpaired Medical Report Generation via Auxiliary Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hirsch, Elad, Dawidowicz, Gefen, Tal, Ayellet
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914879298863104
author Hirsch, Elad
Dawidowicz, Gefen
Tal, Ayellet
author_facet Hirsch, Elad
Dawidowicz, Gefen
Tal, Ayellet
contents Medical report generation from X-ray images is a challenging task, particularly in an unpaired setting where paired image-report data is unavailable for training. To address this challenge, we propose a novel model that leverages the available information in two distinct datasets, one comprising reports and the other consisting of images. The core idea of our model revolves around the notion that combining auto-encoding report generation with multi-modal (report-image) alignment can offer a solution. However, the challenge persists regarding how to achieve this alignment when pair correspondence is absent. Our proposed solution involves the use of auxiliary tasks, particularly contrastive learning and classification, to position related images and reports in close proximity to each other. This approach differs from previous methods that rely on pre-processing steps, such as using external information stored in a knowledge graph. Our model, named MedRAT, surpasses previous state-of-the-art methods, demonstrating the feasibility of generating comprehensive medical reports without the need for paired data or external tools.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03919
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MedRAT: Unpaired Medical Report Generation via Auxiliary Tasks
Hirsch, Elad
Dawidowicz, Gefen
Tal, Ayellet
Computer Vision and Pattern Recognition
Medical report generation from X-ray images is a challenging task, particularly in an unpaired setting where paired image-report data is unavailable for training. To address this challenge, we propose a novel model that leverages the available information in two distinct datasets, one comprising reports and the other consisting of images. The core idea of our model revolves around the notion that combining auto-encoding report generation with multi-modal (report-image) alignment can offer a solution. However, the challenge persists regarding how to achieve this alignment when pair correspondence is absent. Our proposed solution involves the use of auxiliary tasks, particularly contrastive learning and classification, to position related images and reports in close proximity to each other. This approach differs from previous methods that rely on pre-processing steps, such as using external information stored in a knowledge graph. Our model, named MedRAT, surpasses previous state-of-the-art methods, demonstrating the feasibility of generating comprehensive medical reports without the need for paired data or external tools.
title MedRAT: Unpaired Medical Report Generation via Auxiliary Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.03919