Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Haoshuo, Bo, Yufei, Zhang, Hongwei, Tao, Meixia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916927249580032
author Zhang, Haoshuo
Bo, Yufei
Zhang, Hongwei
Tao, Meixia
author_facet Zhang, Haoshuo
Bo, Yufei
Zhang, Hongwei
Tao, Meixia
contents Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the extraction of rich and diverse semantic representations from each modality. To this end, we propose ProMSC-MIS, a Prompt-based Multimodal Semantic Communication system for Multi-spectral Image Segmentation. Specifically, we propose a pre-training algorithm where features from one modality serve as prompts for another, guiding unimodal semantic encoders to learn diverse and complementary semantic representations. We further introduce a semantic fusion module that combines cross-attention mechanisms and squeeze-and-excitation (SE) networks to effectively fuse cross-modal features. Simulation results show that ProMSC-MIS significantly outperforms benchmark methods across various channel-source compression levels, while maintaining low computational complexity and storage overhead. Our scheme has great potential for applications such as autonomous driving and nighttime surveillance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation
Zhang, Haoshuo
Bo, Yufei
Zhang, Hongwei
Tao, Meixia
Image and Video Processing
Multimedia
Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the extraction of rich and diverse semantic representations from each modality. To this end, we propose ProMSC-MIS, a Prompt-based Multimodal Semantic Communication system for Multi-spectral Image Segmentation. Specifically, we propose a pre-training algorithm where features from one modality serve as prompts for another, guiding unimodal semantic encoders to learn diverse and complementary semantic representations. We further introduce a semantic fusion module that combines cross-attention mechanisms and squeeze-and-excitation (SE) networks to effectively fuse cross-modal features. Simulation results show that ProMSC-MIS significantly outperforms benchmark methods across various channel-source compression levels, while maintaining low computational complexity and storage overhead. Our scheme has great potential for applications such as autonomous driving and nighttime surveillance.
title Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation
topic Image and Video Processing
Multimedia
url https://arxiv.org/abs/2508.17920