Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Jingwen, Xiao, Ming, Lyu, Zhonghao, Skoglund, Mikael, Wu, Celimuge
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908645264982016
author Fu, Jingwen
Xiao, Ming
Lyu, Zhonghao
Skoglund, Mikael
Wu, Celimuge
author_facet Fu, Jingwen
Xiao, Ming
Lyu, Zhonghao
Skoglund, Mikael
Wu, Celimuge
contents Semantic communications for multi-modal data can transmit task-relevant information efficiently over noisy and bandwidth-limited channels. However, a key challenge is to simultaneously compress inter-modal redundancy and improve semantic reliability under channel distortion. To address the challenge, we propose a robust and efficient multi-modal task-oriented communication framework that integrates a two-stage variational information bottleneck (VIB) with mutual information (MI) redundancy minimization. In the first stage, we apply uni-modal VIB to compress each modality separately, i.e., text, audio, and video, while preserving task-specific features. To enhance efficiency, an MI minimization module with adversarial training is then used to suppress cross-modal dependencies and to promote complementarity rather than redundancy. In the second stage, a multi-modal VIB is further used to compress the fused representation and to enhance robustness against channel distortion. Experimental results on multi-modal emotion recognition tasks demonstrate that the proposed framework significantly outperforms existing baselines in accuracy and reliability, particularly under low signal-to-noise ratio regimes. Our work provides a principled framework that jointly optimizes modality-specific compression, inter-modal redundancy, and communication reliability.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
Fu, Jingwen
Xiao, Ming
Lyu, Zhonghao
Skoglund, Mikael
Wu, Celimuge
Image and Video Processing
Multimedia
Sound
Semantic communications for multi-modal data can transmit task-relevant information efficiently over noisy and bandwidth-limited channels. However, a key challenge is to simultaneously compress inter-modal redundancy and improve semantic reliability under channel distortion. To address the challenge, we propose a robust and efficient multi-modal task-oriented communication framework that integrates a two-stage variational information bottleneck (VIB) with mutual information (MI) redundancy minimization. In the first stage, we apply uni-modal VIB to compress each modality separately, i.e., text, audio, and video, while preserving task-specific features. To enhance efficiency, an MI minimization module with adversarial training is then used to suppress cross-modal dependencies and to promote complementarity rather than redundancy. In the second stage, a multi-modal VIB is further used to compress the fused representation and to enhance robustness against channel distortion. Experimental results on multi-modal emotion recognition tasks demonstrate that the proposed framework significantly outperforms existing baselines in accuracy and reliability, particularly under low signal-to-noise ratio regimes. Our work provides a principled framework that jointly optimizes modality-specific compression, inter-modal redundancy, and communication reliability.
title Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
topic Image and Video Processing
Multimedia
Sound
url https://arxiv.org/abs/2511.08642