UNIGEOCLIP: Unified Geospatial Contrastive Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Astruc, Guillaume, Trulls, Eduard, Hosang, Jan, Landrieu, Loic, Sarlin, Paul-Edouard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910125647724544
author Astruc, Guillaume
Trulls, Eduard
Hosang, Jan
Landrieu, Loic
Sarlin, Paul-Edouard
author_facet Astruc, Guillaume
Trulls, Eduard
Hosang, Jan
Landrieu, Loic
Sarlin, Paul-Edouard
contents The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11668
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UNIGEOCLIP: Unified Geospatial Contrastive Learning
Astruc, Guillaume
Trulls, Eduard
Hosang, Jan
Landrieu, Loic
Sarlin, Paul-Edouard
Computer Vision and Pattern Recognition
The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.
title UNIGEOCLIP: Unified Geospatial Contrastive Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.11668