GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuhao, Al-Kindi, Sadeer, Veeraraghavan, Ashok, Balakrishnan, Guha
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914565527175168
author Liu, Yuhao
Al-Kindi, Sadeer
Veeraraghavan, Ashok
Balakrishnan, Guha
author_facet Liu, Yuhao
Al-Kindi, Sadeer
Veeraraghavan, Ashok
Balakrishnan, Guha
contents Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured socioeconomic covariates typically stored in tabular form. This modality gap limits their ability to capture the complete total environment, which is critical for reasoning about complex environmental, social, and health-related outcomes. In this work, we propose GeoViSTA (Geospatial Vision-Tabular Transformer), a vision-tabular architecture that learns unified geospatial embeddings from co-registered gridded imagery and tabular data. GeoViSTA utilizes bilateral cross-attention to exchange spatial and semantic information across modalities, guided by a geography-aware attention mechanism that aligns continuous image patches with irregular census-tract tokens. We train GeoViSTA with a self-supervised joint masked-autoencoding objective, forcing it to recover missing image patches and tabular rows using local spatial context and cross-modal cues. Empirically, GeoViSTA's unified embeddings improve linear probing performance on high-impact downstream tasks, outperforming baselines in predicting disease-specific mortality and fire hazard frequency across held-out regions. These results demonstrate that jointly modeling the physical environment alongside structured socioeconomic context yields highly transferable representations for holistic geospatial inference.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14406
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
Liu, Yuhao
Al-Kindi, Sadeer
Veeraraghavan, Ashok
Balakrishnan, Guha
Machine Learning
Computer Vision and Pattern Recognition
Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured socioeconomic covariates typically stored in tabular form. This modality gap limits their ability to capture the complete total environment, which is critical for reasoning about complex environmental, social, and health-related outcomes. In this work, we propose GeoViSTA (Geospatial Vision-Tabular Transformer), a vision-tabular architecture that learns unified geospatial embeddings from co-registered gridded imagery and tabular data. GeoViSTA utilizes bilateral cross-attention to exchange spatial and semantic information across modalities, guided by a geography-aware attention mechanism that aligns continuous image patches with irregular census-tract tokens. We train GeoViSTA with a self-supervised joint masked-autoencoding objective, forcing it to recover missing image patches and tabular rows using local spatial context and cross-modal cues. Empirically, GeoViSTA's unified embeddings improve linear probing performance on high-impact downstream tasks, outperforming baselines in predicting disease-specific mortality and fire hazard frequency across held-out regions. These results demonstrate that jointly modeling the physical environment alongside structured socioeconomic context yields highly transferable representations for holistic geospatial inference.
title GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.14406