Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Yijun, Zhai, Yikui, Ying, Zilu, Xian, Tingfeng, Zhou, Wenlve, Zhou, Zhiheng, Tian, Xiaolin, Jia, Xudong, Zhang, Hongsheng, Chen, C. L. Philip
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911138281684992
author Zhou, Yijun
Zhai, Yikui
Ying, Zilu
Xian, Tingfeng
Zhou, Wenlve
Zhou, Zhiheng
Tian, Xiaolin
Jia, Xudong
Zhang, Hongsheng
Chen, C. L. Philip
author_facet Zhou, Yijun
Zhai, Yikui
Ying, Zilu
Xian, Tingfeng
Zhou, Wenlve
Zhou, Zhiheng
Tian, Xiaolin
Jia, Xudong
Zhang, Hongsheng
Chen, C. L. Philip
contents Although deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision language model (VLM) to generate semantic descriptions of bitemporal images. A Textual Difference Enhancement (TDE) module then captures fine grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image Text Feature Fusion (ITFF) module that enables deep cross modal integration. Extensive experiments on LEVIRCD, WHUCD, and SYSUCD demonstrate that MMChange consistently surpasses state of the art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03961
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
Zhou, Yijun
Zhai, Yikui
Ying, Zilu
Xian, Tingfeng
Zhou, Wenlve
Zhou, Zhiheng
Tian, Xiaolin
Jia, Xudong
Zhang, Hongsheng
Chen, C. L. Philip
Computer Vision and Pattern Recognition
Artificial Intelligence
Although deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision language model (VLM) to generate semantic descriptions of bitemporal images. A Textual Difference Enhancement (TDE) module then captures fine grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image Text Feature Fusion (ITFF) module that enables deep cross modal integration. Extensive experiments on LEVIRCD, WHUCD, and SYSUCD demonstrate that MMChange consistently surpasses state of the art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange.
title Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.03961