BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lozano, Alejandro, Sun, Min Woo, Burgess, James, Chen, Liangyu, Nirschl, Jeffrey J, Gu, Jeffrey, Lopez, Ivan, Aklilu, Josiah, Katzer, Austin Wolfgang, Chiu, Collin, Rau, Anita, Wang, Xiaohan, Zhang, Yuhui, Song, Alfred Seunghoon, Tibshirani, Robert, Yeung-Levy, Serena
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916670724898816
author Lozano, Alejandro
Sun, Min Woo
Burgess, James
Chen, Liangyu
Nirschl, Jeffrey J
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Katzer, Austin Wolfgang
Chiu, Collin
Rau, Anita
Wang, Xiaohan
Zhang, Yuhui
Song, Alfred Seunghoon
Tibshirani, Robert
Yeung-Levy, Serena
author_facet Lozano, Alejandro
Sun, Min Woo
Burgess, James
Chen, Liangyu
Nirschl, Jeffrey J
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Katzer, Austin Wolfgang
Chiu, Collin
Rau, Anita
Wang, Xiaohan
Zhang, Yuhui
Song, Alfred Seunghoon
Tibshirani, Robert
Yeung-Levy, Serena
contents The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are restricted to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA, a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are also provided. We demonstrate the utility and accessibility of our resource by releasing BMCA-CLIP, a suite of CLIP-style models continuously pre-trained on the BIOMEDICA dataset via streaming, eliminating the need to download 27 TB of data locally. On average, our models achieve state-of-the-art performance across 40 tasks - spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology - excelling in zero-shot classification with a 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively), and stronger image-text retrieval, all while using 10x less compute. To foster reproducibility and collaboration, we release our codebase and dataset for the broader research community.
format Preprint
id arxiv_https___arxiv_org_abs_2501_07171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
Lozano, Alejandro
Sun, Min Woo
Burgess, James
Chen, Liangyu
Nirschl, Jeffrey J
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Katzer, Austin Wolfgang
Chiu, Collin
Rau, Anita
Wang, Xiaohan
Zhang, Yuhui
Song, Alfred Seunghoon
Tibshirani, Robert
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Computation and Language
The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are restricted to narrow domains, missing the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA, a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are also provided. We demonstrate the utility and accessibility of our resource by releasing BMCA-CLIP, a suite of CLIP-style models continuously pre-trained on the BIOMEDICA dataset via streaming, eliminating the need to download 27 TB of data locally. On average, our models achieve state-of-the-art performance across 40 tasks - spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology - excelling in zero-shot classification with a 6.56% average improvement (as high as 29.8% and 17.5% in dermatology and ophthalmology, respectively), and stronger image-text retrieval, all while using 10x less compute. To foster reproducibility and collaboration, we release our codebase and dataset for the broader research community.
title BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2501.07171