Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huynh, Andy V., Gillespie, Lauren E., Lopez-Saucedo, Jael, Tang, Claire, Sikand, Rohan, Expósito-Alonso, Moisés
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909328740450304
author Huynh, Andy V.
Gillespie, Lauren E.
Lopez-Saucedo, Jael
Tang, Claire
Sikand, Rohan
Expósito-Alonso, Moisés
author_facet Huynh, Andy V.
Gillespie, Lauren E.
Lopez-Saucedo, Jael
Tang, Claire
Sikand, Rohan
Expósito-Alonso, Moisés
contents Multimodal image-text contrastive learning has shown that joint representations can be learned across modalities. Here, we show how leveraging multiple views of image data with contrastive learning can improve downstream fine-grained classification performance for species recognition, even when one view is absent. We propose ContRastive Image-remote Sensing Pre-training (CRISP)$\unicode{x2014}$a new pre-training task for ground-level and aerial image representation learning of the natural world$\unicode{x2014}$and introduce Nature Multi-View (NMV), a dataset of natural world imagery including $>3$ million ground-level and aerial image pairs for over 6,000 plant taxa across the ecologically diverse state of California. The NMV dataset and accompanying material are available at hf.co/datasets/andyvhuynh/NatureMultiView.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19439
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery
Huynh, Andy V.
Gillespie, Lauren E.
Lopez-Saucedo, Jael
Tang, Claire
Sikand, Rohan
Expósito-Alonso, Moisés
Computer Vision and Pattern Recognition
Multimodal image-text contrastive learning has shown that joint representations can be learned across modalities. Here, we show how leveraging multiple views of image data with contrastive learning can improve downstream fine-grained classification performance for species recognition, even when one view is absent. We propose ContRastive Image-remote Sensing Pre-training (CRISP)$\unicode{x2014}$a new pre-training task for ground-level and aerial image representation learning of the natural world$\unicode{x2014}$and introduce Nature Multi-View (NMV), a dataset of natural world imagery including $>3$ million ground-level and aerial image pairs for over 6,000 plant taxa across the ecologically diverse state of California. The NMV dataset and accompanying material are available at hf.co/datasets/andyvhuynh/NatureMultiView.
title Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.19439