Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909328740450304 |
|---|---|
| author | Huynh, Andy V. Gillespie, Lauren E. Lopez-Saucedo, Jael Tang, Claire Sikand, Rohan Expósito-Alonso, Moisés |
| author_facet | Huynh, Andy V. Gillespie, Lauren E. Lopez-Saucedo, Jael Tang, Claire Sikand, Rohan Expósito-Alonso, Moisés |
| contents | Multimodal image-text contrastive learning has shown that joint representations can be learned across modalities. Here, we show how leveraging multiple views of image data with contrastive learning can improve downstream fine-grained classification performance for species recognition, even when one view is absent. We propose ContRastive Image-remote Sensing Pre-training (CRISP)$\unicode{x2014}$a new pre-training task for ground-level and aerial image representation learning of the natural world$\unicode{x2014}$and introduce Nature Multi-View (NMV), a dataset of natural world imagery including $>3$ million ground-level and aerial image pairs for over 6,000 plant taxa across the ecologically diverse state of California. The NMV dataset and accompanying material are available at hf.co/datasets/andyvhuynh/NatureMultiView. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_19439 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery Huynh, Andy V. Gillespie, Lauren E. Lopez-Saucedo, Jael Tang, Claire Sikand, Rohan Expósito-Alonso, Moisés Computer Vision and Pattern Recognition Multimodal image-text contrastive learning has shown that joint representations can be learned across modalities. Here, we show how leveraging multiple views of image data with contrastive learning can improve downstream fine-grained classification performance for species recognition, even when one view is absent. We propose ContRastive Image-remote Sensing Pre-training (CRISP)$\unicode{x2014}$a new pre-training task for ground-level and aerial image representation learning of the natural world$\unicode{x2014}$and introduce Nature Multi-View (NMV), a dataset of natural world imagery including $>3$ million ground-level and aerial image pairs for over 6,000 plant taxa across the ecologically diverse state of California. The NMV dataset and accompanying material are available at hf.co/datasets/andyvhuynh/NatureMultiView. |
| title | Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.19439 |