Advancing Medical Representation Learning Through High-Quality Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916656103555072 |
|---|---|
| author | Baghbanzadeh, Negin Fallahpour, Adibvafa Parhizkar, Yasaman Ogidi, Franklin Roy, Shuvendu Ashkezari, Sajad Khazaie, Vahid Reza Colacci, Michael Etemad, Ali Afkanpour, Arash Dolatabadi, Elham |
| author_facet | Baghbanzadeh, Negin Fallahpour, Adibvafa Parhizkar, Yasaman Ogidi, Franklin Roy, Shuvendu Ashkezari, Sajad Khazaie, Vahid Reza Colacci, Michael Etemad, Ali Afkanpour, Arash Dolatabadi, Elham |
| contents | Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, containing 2.2 million image-text pairs, enriched with image modality annotations, subfigures, and summarized in-text references. Notably, the in-text references provide richer medical context, extending beyond the abstract information typically found in captions. Through extensive experiments, we benchmark Open-PMC against larger datasets across retrieval and zero-shot classification tasks. Our results show that dataset quality-not just size-drives significant performance gains. We complement our benchmark with an in-depth analysis of feature representation. Our findings highlight the crucial role of data curation quality in advancing multimodal medical AI. We release Open-PMC, along with the trained models and our codebase. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_14377 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Advancing Medical Representation Learning Through High-Quality Data Baghbanzadeh, Negin Fallahpour, Adibvafa Parhizkar, Yasaman Ogidi, Franklin Roy, Shuvendu Ashkezari, Sajad Khazaie, Vahid Reza Colacci, Michael Etemad, Ali Afkanpour, Arash Dolatabadi, Elham Image and Video Processing Computer Vision and Pattern Recognition Machine Learning Despite the growing scale of medical Vision-Language datasets, the impact of dataset quality on model performance remains under-explored. We introduce Open-PMC, a high-quality medical dataset from PubMed Central, containing 2.2 million image-text pairs, enriched with image modality annotations, subfigures, and summarized in-text references. Notably, the in-text references provide richer medical context, extending beyond the abstract information typically found in captions. Through extensive experiments, we benchmark Open-PMC against larger datasets across retrieval and zero-shot classification tasks. Our results show that dataset quality-not just size-drives significant performance gains. We complement our benchmark with an in-depth analysis of feature representation. Our findings highlight the crucial role of data curation quality in advancing multimodal medical AI. We release Open-PMC, along with the trained models and our codebase. |
| title | Advancing Medical Representation Learning Through High-Quality Data |
| topic | Image and Video Processing Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2503.14377 |