SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Futian, Wang, Mengqi, Wang, Xiao, Wang, Haowen, Tang, Jin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917106568658944
author Wang, Futian
Wang, Mengqi
Wang, Xiao
Wang, Haowen
Tang, Jin
author_facet Wang, Futian
Wang, Mengqi
Wang, Xiao
Wang, Haowen
Tang, Jin
contents Remote sensing change captioning is an emerging and popular research task that aims to describe, in natural language, the content of interest that has changed between two remote sensing images captured at different times. Existing methods typically employ CNNs/Transformers to extract visual representations from the given images or incorporate auxiliary tasks to enhance the final results, with weak region awareness and limited temporal alignment. To address these issues, this paper explores the use of the SAM (Segment Anything Model) foundation model to extract region-level representations and inject region-of-interest knowledge into the captioning framework. Specifically, we employ a CNN/Transformer model to extract global-level vision features, leverage the SAM foundation model to delineate semantic- and motion-level change regions, and utilize a specially constructed knowledge graph to provide information about objects of interest. These heterogeneous sources of information are then fused via cross-attention, and a Transformer decoder is used to generate the final natural language description of the observed changes. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple widely used benchmark datasets. The source code of this paper will be released on https://github.com/Event-AHU/SAM_ChangeCaptioning
format Preprint
id arxiv_https___arxiv_org_abs_2511_21420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
Wang, Futian
Wang, Mengqi
Wang, Xiao
Wang, Haowen
Tang, Jin
Computer Vision and Pattern Recognition
Artificial Intelligence
Remote sensing change captioning is an emerging and popular research task that aims to describe, in natural language, the content of interest that has changed between two remote sensing images captured at different times. Existing methods typically employ CNNs/Transformers to extract visual representations from the given images or incorporate auxiliary tasks to enhance the final results, with weak region awareness and limited temporal alignment. To address these issues, this paper explores the use of the SAM (Segment Anything Model) foundation model to extract region-level representations and inject region-of-interest knowledge into the captioning framework. Specifically, we employ a CNN/Transformer model to extract global-level vision features, leverage the SAM foundation model to delineate semantic- and motion-level change regions, and utilize a specially constructed knowledge graph to provide information about objects of interest. These heterogeneous sources of information are then fused via cross-attention, and a Transformer decoder is used to generate the final natural language description of the observed changes. Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple widely used benchmark datasets. The source code of this paper will be released on https://github.com/Event-AHU/SAM_ChangeCaptioning
title SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.21420