SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Dongchen, Wang, Di, Gao, Erzhong, Qin, Xiaolei, Zhao, Liu, Zhang, Jing, Xu, Minqiang, Zhan, Jianbo, Wang, Jianshe, Liu, Lin, Du, Bo, Zhang, Liangpei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914379081973760
author Si, Dongchen
Wang, Di
Gao, Erzhong
Qin, Xiaolei
Zhao, Liu
Zhang, Jing
Xu, Minqiang
Zhan, Jianbo
Wang, Jianshe
Liu, Lin
Du, Bo
Zhang, Liangpei
author_facet Si, Dongchen
Wang, Di
Gao, Erzhong
Qin, Xiaolei
Zhao, Liu
Zhang, Jing
Xu, Minqiang
Zhan, Jianbo
Wang, Jianshe
Liu, Lin
Du, Bo
Zhang, Liangpei
contents Spectral information has long been recognized as a critical cue in remote sensing observations. Although numerous vision-language models have been developed for pixel-level interpretation, spectral information remains underutilized, resulting in suboptimal performance, particularly in multispectral scenarios. To address this limitation, we construct a vision-language instruction-following dataset named SPIE, which encodes spectral priors of land-cover objects into textual attributes recognizable by large language models (LLMs), based on classical spectral index computations. Leveraging this dataset, we propose SPEX, a multimodal LLM designed for instruction-driven land cover extraction. To this end, we introduce several carefully designed components and training strategies, including multiscale feature aggregation, token context condensation, and multispectral visual pre-training, to achieve precise and flexible pixel-level interpretation. To the best of our knowledge, SPEX is the first multimodal vision-language model dedicated to land cover extraction in spectral remote sensing imagery. Extensive experiments on five public multispectral datasets demonstrate that SPEX consistently outperforms existing state-of-the-art methods in extracting typical land cover categories such as vegetation, buildings, and water bodies. Moreover, SPEX is capable of generating textual explanations for its predictions, thereby enhancing interpretability and user-friendliness. Code will be released at: https://github.com/MiliLab/SPEX.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
Si, Dongchen
Wang, Di
Gao, Erzhong
Qin, Xiaolei
Zhao, Liu
Zhang, Jing
Xu, Minqiang
Zhan, Jianbo
Wang, Jianshe
Liu, Lin
Du, Bo
Zhang, Liangpei
Computer Vision and Pattern Recognition
Spectral information has long been recognized as a critical cue in remote sensing observations. Although numerous vision-language models have been developed for pixel-level interpretation, spectral information remains underutilized, resulting in suboptimal performance, particularly in multispectral scenarios. To address this limitation, we construct a vision-language instruction-following dataset named SPIE, which encodes spectral priors of land-cover objects into textual attributes recognizable by large language models (LLMs), based on classical spectral index computations. Leveraging this dataset, we propose SPEX, a multimodal LLM designed for instruction-driven land cover extraction. To this end, we introduce several carefully designed components and training strategies, including multiscale feature aggregation, token context condensation, and multispectral visual pre-training, to achieve precise and flexible pixel-level interpretation. To the best of our knowledge, SPEX is the first multimodal vision-language model dedicated to land cover extraction in spectral remote sensing imagery. Extensive experiments on five public multispectral datasets demonstrate that SPEX consistently outperforms existing state-of-the-art methods in extracting typical land cover categories such as vegetation, buildings, and water bodies. Moreover, SPEX is capable of generating textual explanations for its predictions, thereby enhancing interpretability and user-friendliness. Code will be released at: https://github.com/MiliLab/SPEX.
title SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05202