BioCLIP: A Vision Foundation Model for the Tree of Life

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stevens, Samuel, Wu, Jiaman, Thompson, Matthew J, Campolongo, Elizabeth G, Song, Chan Hee, Carlyn, David Edward, Dong, Li, Dahdul, Wasila M, Stewart, Charles, Berger-Wolf, Tanya, Chao, Wei-Lun, Su, Yu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909202590466048
author Stevens, Samuel
Wu, Jiaman
Thompson, Matthew J
Campolongo, Elizabeth G
Song, Chan Hee
Carlyn, David Edward
Dong, Li
Dahdul, Wasila M
Stewart, Charles
Berger-Wolf, Tanya
Chao, Wei-Lun
Su, Yu
author_facet Stevens, Samuel
Wu, Jiaman
Thompson, Matthew J
Campolongo, Elizabeth G
Song, Chan Hee
Carlyn, David Edward
Dong, Li
Dahdul, Wasila M
Stewart, Charles
Berger-Wolf, Tanya
Chao, Wei-Lun
Su, Yu
contents Images of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an explosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general organismal biology questions on images is of timely need. To approach this, we curate and release TreeOfLife-10M, the largest and most diverse ML-ready dataset of biology images. We then develop BioCLIP, a foundation model for the tree of life, leveraging the unique properties of biology captured by TreeOfLife-10M, namely the abundance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on diverse fine-grained biology classification tasks and find that BioCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BioCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability. https://imageomics.github.io/bioclip has models, data and code.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18803
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle BioCLIP: A Vision Foundation Model for the Tree of Life
Stevens, Samuel
Wu, Jiaman
Thompson, Matthew J
Campolongo, Elizabeth G
Song, Chan Hee
Carlyn, David Edward
Dong, Li
Dahdul, Wasila M
Stewart, Charles
Berger-Wolf, Tanya
Chao, Wei-Lun
Su, Yu
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Images of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an explosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general organismal biology questions on images is of timely need. To approach this, we curate and release TreeOfLife-10M, the largest and most diverse ML-ready dataset of biology images. We then develop BioCLIP, a foundation model for the tree of life, leveraging the unique properties of biology captured by TreeOfLife-10M, namely the abundance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on diverse fine-grained biology classification tasks and find that BioCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BioCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability. https://imageomics.github.io/bioclip has models, data and code.
title BioCLIP: A Vision Foundation Model for the Tree of Life
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2311.18803