TaxaBind: A Unified Embedding Space for Ecological Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sastry, Srikumar, Khanal, Subash, Dhakal, Aayush, Ahmad, Adeel, Jacobs, Nathan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909374658641920
author Sastry, Srikumar
Khanal, Subash
Dhakal, Aayush
Ahmad, Adeel
Jacobs, Nathan
author_facet Sastry, Srikumar
Khanal, Subash
Dhakal, Aayush
Ahmad, Adeel
Jacobs, Nathan
contents We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving ecological problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emergent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvrl/TaxaBind.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00683
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TaxaBind: A Unified Embedding Space for Ecological Applications
Sastry, Srikumar
Khanal, Subash
Dhakal, Aayush
Ahmad, Adeel
Jacobs, Nathan
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving ecological problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emergent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvrl/TaxaBind.
title TaxaBind: A Unified Embedding Space for Ecological Applications
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.00683