Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ko, Hanbin, Cho, Gihun, Baek, Inhyeok, Kim, Donguk, Koo, Joonbeom, Kim, Changi, Lee, Dongheon, Park, Chang Min
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918534135676928
author Ko, Hanbin
Cho, Gihun
Baek, Inhyeok
Kim, Donguk
Koo, Joonbeom
Kim, Changi
Lee, Dongheon
Park, Chang Min
author_facet Ko, Hanbin
Cho, Gihun
Baek, Inhyeok
Kim, Donguk
Koo, Joonbeom
Kim, Changi
Lee, Dongheon
Park, Chang Min
contents Multimodal learning from paired medical images and clinical text is a central challenge in medical data-driven informatics, where effective cross-modal alignment is critical for scalable analysis and retrieval. In chest radiography, vision-language pretraining is constrained by heterogeneous radiology reports that contain abbreviations, impression-only notes, and institution-specific writing styles. Unlike general-domain settings, naively aggregating large collections of noisy reports can plateau or even degrade multimodal learning when reporting styles differ substantially. We propose a domain-adapted bidirectional large language model text encoder for chest radiograph reports, trained with masked token prediction and supervised contrastive learning on stylistically diverse but clinically equivalent report variants to produce robust, generalizable text embeddings. We then integrate this encoder into a dual-tower contrastive vision-language framework using parameter-efficient adaptation to improve image-text alignment. Across 1.6 million paired studies from public datasets and a de-identified hospital cohort, the proposed models improve bidirectional retrieval accuracy and external generalization, achieving GREEN scores of 0.308 on MIMIC-CXR and 0.618 on Open-I, while reducing the degradation observed when abbreviation-rich, impression-only hospital reports are added to training.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
Ko, Hanbin
Cho, Gihun
Baek, Inhyeok
Kim, Donguk
Koo, Joonbeom
Kim, Changi
Lee, Dongheon
Park, Chang Min
Computer Vision and Pattern Recognition
68T07, 68U10, 92C55
I.2.10; I.2.7
Multimodal learning from paired medical images and clinical text is a central challenge in medical data-driven informatics, where effective cross-modal alignment is critical for scalable analysis and retrieval. In chest radiography, vision-language pretraining is constrained by heterogeneous radiology reports that contain abbreviations, impression-only notes, and institution-specific writing styles. Unlike general-domain settings, naively aggregating large collections of noisy reports can plateau or even degrade multimodal learning when reporting styles differ substantially. We propose a domain-adapted bidirectional large language model text encoder for chest radiograph reports, trained with masked token prediction and supervised contrastive learning on stylistically diverse but clinically equivalent report variants to produce robust, generalizable text embeddings. We then integrate this encoder into a dual-tower contrastive vision-language framework using parameter-efficient adaptation to improve image-text alignment. Across 1.6 million paired studies from public datasets and a de-identified hospital cohort, the proposed models improve bidirectional retrieval accuracy and external generalization, achieving GREEN scores of 0.308 on MIMIC-CXR and 0.618 on Open-I, while reducing the degradation observed when abbreviation-rich, impression-only hospital reports are added to training.
title Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
topic Computer Vision and Pattern Recognition
68T07, 68U10, 92C55
I.2.10; I.2.7
url https://arxiv.org/abs/2509.15234