Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.22727 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912305167466496 |
|---|---|
| author | Lozano, Alejandro Sun, Min Woo Burgess, James Nirschl, Jeffrey J. Polzak, Christopher Zhang, Yuhui Chen, Liangyu Gu, Jeffrey Lopez, Ivan Aklilu, Josiah Rau, Anita Katzer, Austin Wolfgang Chiu, Collin Zohar, Orr Wang, Xiaohan Song, Alfred Seunghoon Chia-Chun, Chiang Tibshirani, Robert Yeung-Levy, Serena |
| author_facet | Lozano, Alejandro Sun, Min Woo Burgess, James Nirschl, Jeffrey J. Polzak, Christopher Zhang, Yuhui Chen, Liangyu Gu, Jeffrey Lopez, Ivan Aklilu, Josiah Rau, Anita Katzer, Austin Wolfgang Chiu, Collin Zohar, Orr Wang, Xiaohan Song, Alfred Seunghoon Chia-Chun, Chiang Tibshirani, Robert Yeung-Levy, Serena |
| contents | Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_22727 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI Lozano, Alejandro Sun, Min Woo Burgess, James Nirschl, Jeffrey J. Polzak, Christopher Zhang, Yuhui Chen, Liangyu Gu, Jeffrey Lopez, Ivan Aklilu, Josiah Rau, Anita Katzer, Austin Wolfgang Chiu, Collin Zohar, Orr Wang, Xiaohan Song, Alfred Seunghoon Chia-Chun, Chiang Tibshirani, Robert Yeung-Levy, Serena Computation and Language Machine Learning Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 million scientific articles and 24 million image-text pairs, along with 27 metadata fields (including expert human annotations). To overcome the challenges of accessing our large-scale dataset, we provide scalable streaming and search APIs through a web server, facilitating seamless integration with AI systems. We demonstrate the utility of the Biomedica dataset by building embedding models, chat-style models, and retrieval-augmented chat agents. Notably, all our AI models surpass previous open systems in their respective categories, underscoring the critical role of diverse, high-quality, and large-scale biomedical data. |
| title | A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2503.22727 |