TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shu, Yan, Ren, Bin, Xiong, Zhitong, Zhu, Xiao Xiang, Demir, Begüm, Sebe, Nicu, Rota, Paolo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912974469332992
author Shu, Yan
Ren, Bin
Xiong, Zhitong
Zhu, Xiao Xiang
Demir, Begüm
Sebe, Nicu
Rota, Paolo
author_facet Shu, Yan
Ren, Bin
Xiong, Zhitong
Zhu, Xiao Xiang
Demir, Begüm
Sebe, Nicu
Rota, Paolo
contents Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospatial reasoning with two key capabilities: (1) modality-flexible reasoning: it handles single-modality inputs (optical or SAR) and adaptively fuses different modalities into the reasoning process when both are available; (2) multi-temporal reasoning: it integrates temporal sequences for change analysis across multiple time points. In addition, we curate Terra-CoT, a large-scale dataset containing 1 million samples with pixel-level masks embedded in reasoning chains across multiple sources. We also propose TerraScope-Bench, the first benchmark for pixel-grounded geospatial reasoning with six sub-tasks that evaluates both answer accuracy and mask quality to ensure authentic pixel-grounded reasoning. Experiments show that TerraScope significantly outperforms existing VLMs on pixel-grounded geospatial reasoning while providing interpretable visual evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19039
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
Shu, Yan
Ren, Bin
Xiong, Zhitong
Zhu, Xiao Xiang
Demir, Begüm
Sebe, Nicu
Rota, Paolo
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospatial reasoning with two key capabilities: (1) modality-flexible reasoning: it handles single-modality inputs (optical or SAR) and adaptively fuses different modalities into the reasoning process when both are available; (2) multi-temporal reasoning: it integrates temporal sequences for change analysis across multiple time points. In addition, we curate Terra-CoT, a large-scale dataset containing 1 million samples with pixel-level masks embedded in reasoning chains across multiple sources. We also propose TerraScope-Bench, the first benchmark for pixel-grounded geospatial reasoning with six sub-tasks that evaluates both answer accuracy and mask quality to ensure authentic pixel-grounded reasoning. Experiments show that TerraScope significantly outperforms existing VLMs on pixel-grounded geospatial reasoning while providing interpretable visual evidence.
title TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.19039