Advancing Medical Representation Learning Through High-Quality Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Baghbanzadeh, Negin, Fallahpour, Adibvafa, Parhizkar, Yasaman, Ogidi, Franklin, Roy, Shuvendu, Ashkezari, Sajad, Khazaie, Vahid Reza, Colacci, Michael, Etemad, Ali, Afkanpour, Arash, Dolatabadi, Elham
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916656103555072
author Baghbanzadeh, Negin
Fallahpour, Adibvafa
Parhizkar, Yasaman
Ogidi, Franklin
Roy, Shuvendu
Ashkezari, Sajad
Khazaie, Vahid Reza
Colacci, Michael
Etemad, Ali
Afkanpour, Arash
Dolatabadi, Elham
author_facet Baghbanzadeh, Negin
Fallahpour, Adibvafa
Parhizkar, Yasaman
Ogidi, Franklin
Roy, Shuvendu
Ashkezari, Sajad
Khazaie, Vahid Reza
Colacci, Michael
Etemad, Ali
Afkanpour, Arash
Dolatabadi, Elham
contents Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, containing 2.2 million image-text pairs, enriched with image modality annotations, subfigures, and summarized in-text references. Notably, the in-text references provide richer medical context, extending beyond the abstract information typically found in captions. Through extensive experiments, we benchmark Open-PMC against larger datasets across retrieval and zero-shot classification tasks. Our results show that dataset quality-not just size-drives significant performance gains. We complement our benchmark with an in-depth analysis of feature representation. Our findings highlight the crucial role of data curation quality in advancing multimodal medical AI. We release Open-PMC, along with the trained models and our codebase.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Advancing Medical Representation Learning Through High-Quality Data
Baghbanzadeh, Negin
Fallahpour, Adibvafa
Parhizkar, Yasaman
Ogidi, Franklin
Roy, Shuvendu
Ashkezari, Sajad
Khazaie, Vahid Reza
Colacci, Michael
Etemad, Ali
Afkanpour, Arash
Dolatabadi, Elham
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, containing 2.2 million image-text pairs, enriched with image modality annotations, subfigures, and summarized in-text references. Notably, the in-text references provide richer medical context, extending beyond the abstract information typically found in captions. Through extensive experiments, we benchmark Open-PMC against larger datasets across retrieval and zero-shot classification tasks. Our results show that dataset quality-not just size-drives significant performance gains. We complement our benchmark with an in-depth analysis of feature representation. Our findings highlight the crucial role of data curation quality in advancing multimodal medical AI. We release Open-PMC, along with the trained models and our codebase.
title Advancing Medical Representation Learning Through High-Quality Data
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.14377