GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hu, Xuran, Xiong, Zhitong, Hong, Zhongcheng, Ban, Yifang, Zhu, Xiaoxiang, Zhao, Wufan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918411178606592
author Hu, Xuran
Xiong, Zhitong
Hong, Zhongcheng
Ban, Yifang
Zhu, Xiaoxiang
Zhao, Wufan
author_facet Hu, Xuran
Xiong, Zhitong
Hong, Zhongcheng
Ban, Yifang
Zhu, Xiaoxiang
Zhao, Wufan
contents Current Large Multimodal Models (LMMs) in Earth Observation typically neglect the critical "vertical" dimension, limiting their reasoning capabilities in complex remote sensing geometries and disaster scenarios where physical spatial structures often outweigh planar visual textures. To bridge this gap, we introduce a comprehensive evaluation framework dedicated to height-aware remote sensing understanding. First, to overcome the severe scarcity of annotated data, we develop a scalable, VLM-driven data generation pipeline utilizing systematic prompt engineering and metadata extraction. This pipeline constructs two complementary benchmarks: GeoHeight-Bench for relative height analysis, and a more challenging GeoHeight-Bench+ for holistic, terrain-aware reasoning. Furthermore, to validate the necessity of height perception, we propose GeoHeightChat, the first height-aware remote sensing LMM baseline. Serving as a strong proof of concept, our baseline demonstrates that synergizing visual semantics with implicitly injected height geometric features effectively mitigates the "vertical blind spot", successfully unlocking a new paradigm of interactive height reasoning in existing optical models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25565
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
Hu, Xuran
Xiong, Zhitong
Hong, Zhongcheng
Ban, Yifang
Zhu, Xiaoxiang
Zhao, Wufan
Computer Vision and Pattern Recognition
I.2.10
Current Large Multimodal Models (LMMs) in Earth Observation typically neglect the critical "vertical" dimension, limiting their reasoning capabilities in complex remote sensing geometries and disaster scenarios where physical spatial structures often outweigh planar visual textures. To bridge this gap, we introduce a comprehensive evaluation framework dedicated to height-aware remote sensing understanding. First, to overcome the severe scarcity of annotated data, we develop a scalable, VLM-driven data generation pipeline utilizing systematic prompt engineering and metadata extraction. This pipeline constructs two complementary benchmarks: GeoHeight-Bench for relative height analysis, and a more challenging GeoHeight-Bench+ for holistic, terrain-aware reasoning. Furthermore, to validate the necessity of height perception, we propose GeoHeightChat, the first height-aware remote sensing LMM baseline. Serving as a strong proof of concept, our baseline demonstrates that synergizing visual semantics with implicitly injected height geometric features effectively mitigates the "vertical blind spot", successfully unlocking a new paradigm of interactive height reasoning in existing optical models.
title GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
topic Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2603.25565