UNIGEOCLIP: Unified Geospatial Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910125647724544 |
|---|---|
| author | Astruc, Guillaume Trulls, Eduard Hosang, Jan Landrieu, Loic Sarlin, Paul-Edouard |
| author_facet | Astruc, Guillaume Trulls, Eduard Hosang, Jan Landrieu, Loic Sarlin, Paul-Edouard |
| contents | The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_11668 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | UNIGEOCLIP: Unified Geospatial Contrastive Learning Astruc, Guillaume Trulls, Eduard Hosang, Jan Landrieu, Loic Sarlin, Paul-Edouard Computer Vision and Pattern Recognition The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip. |
| title | UNIGEOCLIP: Unified Geospatial Contrastive Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.11668 |