A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929684426522624 |
|---|---|
| author | Ma, Xiaowen Lian, Rongrong Wu, Zhenkai Guan, Renxiang Hong, Tingfeng Zhao, Mengjiao Ma, Mengting Nie, Jiangtao Du, Zhenhong Song, Siyang Zhang, Wei |
| author_facet | Ma, Xiaowen Lian, Rongrong Wu, Zhenkai Guan, Renxiang Hong, Tingfeng Zhao, Mengjiao Ma, Mengting Nie, Jiangtao Du, Zhenhong Song, Siyang Zhang, Wei |
| contents | As a common method in the field of computer vision, spatial attention mechanism has been widely used in semantic segmentation of remote sensing images due to its outstanding long-range dependency modeling capability. However, remote sensing images are usually characterized by complex backgrounds and large intra-class variance that would degrade their analysis performance. While vanilla spatial attention mechanisms are based on dense affine operations, they tend to introduce a large amount of background contextual information and lack of consideration for intrinsic spatial correlation. To deal with such limitations, this paper proposes a novel scene-Coupling semantic mask network, which reconstructs the vanilla attention with scene coupling and local global semantic masks strategies. Specifically, scene coupling module decomposes scene information into global representations and object distributions, which are then embedded in the attention affinity processes. This Strategy effectively utilizes the intrinsic spatial correlation between features so that improve the process of attention modeling. Meanwhile, local global semantic masks module indirectly correlate pixels with the global semantic masks by using the local semantic mask as an intermediate sensory element, which reduces the background contextual interference and mitigates the effect of intra-class variance. By combining the above two strategies, we propose the model SCSM, which not only can efficiently segment various geospatial objects in complex scenarios, but also possesses inter-clean and elegant mathematical representations. Experimental results on four benchmark datasets demonstrate the the effectiveness of the above two strategies for improving the attention modeling of remote sensing images. The dataset and code are available at https://github.com/xwmaxwma/rssegmentation |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_13130 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation Ma, Xiaowen Lian, Rongrong Wu, Zhenkai Guan, Renxiang Hong, Tingfeng Zhao, Mengjiao Ma, Mengting Nie, Jiangtao Du, Zhenhong Song, Siyang Zhang, Wei Image and Video Processing As a common method in the field of computer vision, spatial attention mechanism has been widely used in semantic segmentation of remote sensing images due to its outstanding long-range dependency modeling capability. However, remote sensing images are usually characterized by complex backgrounds and large intra-class variance that would degrade their analysis performance. While vanilla spatial attention mechanisms are based on dense affine operations, they tend to introduce a large amount of background contextual information and lack of consideration for intrinsic spatial correlation. To deal with such limitations, this paper proposes a novel scene-Coupling semantic mask network, which reconstructs the vanilla attention with scene coupling and local global semantic masks strategies. Specifically, scene coupling module decomposes scene information into global representations and object distributions, which are then embedded in the attention affinity processes. This Strategy effectively utilizes the intrinsic spatial correlation between features so that improve the process of attention modeling. Meanwhile, local global semantic masks module indirectly correlate pixels with the global semantic masks by using the local semantic mask as an intermediate sensory element, which reduces the background contextual interference and mitigates the effect of intra-class variance. By combining the above two strategies, we propose the model SCSM, which not only can efficiently segment various geospatial objects in complex scenarios, but also possesses inter-clean and elegant mathematical representations. Experimental results on four benchmark datasets demonstrate the the effectiveness of the above two strategies for improving the attention modeling of remote sensing images. The dataset and code are available at https://github.com/xwmaxwma/rssegmentation |
| title | A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation |
| topic | Image and Video Processing |
| url | https://arxiv.org/abs/2501.13130 |