TaxaBind: A Unified Embedding Space for Ecological Applications
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909374658641920 |
|---|---|
| author | Sastry, Srikumar Khanal, Subash Dhakal, Aayush Ahmad, Adeel Jacobs, Nathan |
| author_facet | Sastry, Srikumar Khanal, Subash Dhakal, Aayush Ahmad, Adeel Jacobs, Nathan |
| contents | We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving ecological problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emergent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvrl/TaxaBind. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_00683 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | TaxaBind: A Unified Embedding Space for Ecological Applications Sastry, Srikumar Khanal, Subash Dhakal, Aayush Ahmad, Adeel Jacobs, Nathan Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving ecological problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emergent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvrl/TaxaBind. |
| title | TaxaBind: A Unified Embedding Space for Ecological Applications |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2411.00683 |