HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xinyu, Mai, Zurong, Li, Qingmei, Liao, Zjin, Wen, Yibin, Chen, Yuhang, Fan, Xiaoya, Ho, Chan Tsz, Tianyuan, Bi, Liang, Haoyuan, Su, Ruifeng, Qian, Zihao, Zheng, Juepeng, Huang, Jianxi, Lu, Yutong, Fu, Haohuan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917398449225728
author Zhang, Xinyu
Mai, Zurong
Li, Qingmei
Liao, Zjin
Wen, Yibin
Chen, Yuhang
Fan, Xiaoya
Ho, Chan Tsz
Tianyuan, Bi
Liang, Haoyuan
Su, Ruifeng
Qian, Zihao
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
author_facet Zhang, Xinyu
Mai, Zurong
Li, Qingmei
Liao, Zjin
Wen, Yibin
Chen, Yuhang
Fan, Xiaoya
Ho, Chan Tsz
Tianyuan, Bi
Liang, Haoyuan
Su, Ruifeng
Qian, Zihao
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
contents While multimodal large language models (MLLMs) have made significant strides in natural image understanding, their ability to perceive and reason over hyperspectral image (HSI) remains underexplored, which is a vital modality in remote sensing. The high dimensionality and intricate spectral-spatial properties of HSI pose unique challenges for models primarily trained on RGB data.To address this gap, we introduce Hyperspectral Multimodal Benchmark (HM-Bench), the first benchmark designed specifically to evaluate MLLMs in HSI understanding. We curate a large-scale dataset of 19,337 question-answer pairs across 13 task categories, ranging from basic perception to spectral reasoning. Given that existing MLLMs are not equipped to process raw hyperspectral cubes natively, we propose a dual-modality evaluation framework that transforms HSI data into two complementary representations: PCA-based composite images and structured textual reports. This approach facilitates a systematic comparison of different representation for model performance. Extensive evaluations on 18 representative MLLMs reveal significant difficulties in handling complex spatial-spectral reasoning tasks. Furthermore, our results demonstrate that visual inputs generally outperform textual inputs, highlighting the importance of grounding in spectral-spatial evidence for effective HSI understanding. Dataset and appendix can be accessed at https://github.com/HuoRiLi-Yu/HM-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08884
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
Zhang, Xinyu
Mai, Zurong
Li, Qingmei
Liao, Zjin
Wen, Yibin
Chen, Yuhang
Fan, Xiaoya
Ho, Chan Tsz
Tianyuan, Bi
Liang, Haoyuan
Su, Ruifeng
Qian, Zihao
Zheng, Juepeng
Huang, Jianxi
Lu, Yutong
Fu, Haohuan
Computer Vision and Pattern Recognition
Artificial Intelligence
While multimodal large language models (MLLMs) have made significant strides in natural image understanding, their ability to perceive and reason over hyperspectral image (HSI) remains underexplored, which is a vital modality in remote sensing. The high dimensionality and intricate spectral-spatial properties of HSI pose unique challenges for models primarily trained on RGB data.To address this gap, we introduce Hyperspectral Multimodal Benchmark (HM-Bench), the first benchmark designed specifically to evaluate MLLMs in HSI understanding. We curate a large-scale dataset of 19,337 question-answer pairs across 13 task categories, ranging from basic perception to spectral reasoning. Given that existing MLLMs are not equipped to process raw hyperspectral cubes natively, we propose a dual-modality evaluation framework that transforms HSI data into two complementary representations: PCA-based composite images and structured textual reports. This approach facilitates a systematic comparison of different representation for model performance. Extensive evaluations on 18 representative MLLMs reveal significant difficulties in handling complex spatial-spectral reasoning tasks. Furthermore, our results demonstrate that visual inputs generally outperform textual inputs, highlighting the importance of grounding in spectral-spatial evidence for effective HSI understanding. Dataset and appendix can be accessed at https://github.com/HuoRiLi-Yu/HM-Bench.
title HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.08884