Saved in:
Bibliographic Details
Main Authors: Lozano, Alejandro, Sun, Min Woo, Burgess, James, Nirschl, Jeffrey J., Polzak, Christopher, Zhang, Yuhui, Chen, Liangyu, Gu, Jeffrey, Lopez, Ivan, Aklilu, Josiah, Rau, Anita, Katzer, Austin Wolfgang, Chiu, Collin, Zohar, Orr, Wang, Xiaohan, Song, Alfred Seunghoon, Chia-Chun, Chiang, Tibshirani, Robert, Yeung-Levy, Serena
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.22727
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912305167466496
author Lozano, Alejandro
Sun, Min Woo
Burgess, James
Nirschl, Jeffrey J.
Polzak, Christopher
Zhang, Yuhui
Chen, Liangyu
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Rau, Anita
Katzer, Austin Wolfgang
Chiu, Collin
Zohar, Orr
Wang, Xiaohan
Song, Alfred Seunghoon
Chia-Chun, Chiang
Tibshirani, Robert
Yeung-Levy, Serena
author_facet Lozano, Alejandro
Sun, Min Woo
Burgess, James
Nirschl, Jeffrey J.
Polzak, Christopher
Zhang, Yuhui
Chen, Liangyu
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Rau, Anita
Katzer, Austin Wolfgang
Chiu, Collin
Zohar, Orr
Wang, Xiaohan
Song, Alfred Seunghoon
Chia-Chun, Chiang
Tibshirani, Robert
Yeung-Levy, Serena
contents Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22727
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
Lozano, Alejandro
Sun, Min Woo
Burgess, James
Nirschl, Jeffrey J.
Polzak, Christopher
Zhang, Yuhui
Chen, Liangyu
Gu, Jeffrey
Lopez, Ivan
Aklilu, Josiah
Rau, Anita
Katzer, Austin Wolfgang
Chiu, Collin
Zohar, Orr
Wang, Xiaohan
Song, Alfred Seunghoon
Chia-Chun, Chiang
Tibshirani, Robert
Yeung-Levy, Serena
Computation and Language
Machine Learning
Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data.
title A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.22727