Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Jingwen, Xiao, Ming, Lyu, Zhonghao, Skoglund, Mikael, Wu, Celimuge |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Computation-resource-efficient Task-oriented Communications
by: Fu, Jingwen, et al.
Published: (2025)
by: Fu, Jingwen, et al.
Published: (2025)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024)
by: Yue, Xianghu, et al.
Published: (2024)
DriftDecode: One-Step Wireless Image Decoding via Drifting-Inspired Detail Recovery
by: Fu, Jingwen, et al.
Published: (2026)
by: Fu, Jingwen, et al.
Published: (2026)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
by: Ragano, Alessandro, et al.
Published: (2024)
by: Ragano, Alessandro, et al.
Published: (2024)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)
by: Mingote, Victoria, et al.
Published: (2024)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Is there a relationship between Mean Opinion Score (MOS) and Just Noticeable Difference (JND)?
by: Zhu, Jingwen, et al.
Published: (2026)
by: Zhu, Jingwen, et al.
Published: (2026)
Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation
by: Zhang, Haoshuo, et al.
Published: (2025)
by: Zhang, Haoshuo, et al.
Published: (2025)
Robust Live Streaming over LEO Satellite Constellations: Measurement, Analysis, and Handover-Aware Adaptation
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Land-then-transport: A Flow Matching-Based Generative Decoder for Wireless Image Transmission
by: Fu, Jingwen, et al.
Published: (2026)
by: Fu, Jingwen, et al.
Published: (2026)
Perception-Aware Video Semantic Communication
by: Huang, Yinhuan, et al.
Published: (2026)
by: Huang, Yinhuan, et al.
Published: (2026)
Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective
by: Chen, Peilin, et al.
Published: (2024)
by: Chen, Peilin, et al.
Published: (2024)
Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
by: Jin, Yili, et al.
Published: (2025)
by: Jin, Yili, et al.
Published: (2025)
Multi-hop Parallel Image Semantic Communication for Distortion Accumulation Mitigation
by: Xie, Bingyan, et al.
Published: (2025)
by: Xie, Bingyan, et al.
Published: (2025)
H.265/HEVC Video Steganalysis Based on CU Block Structure Gradients and IPM Mapping
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
QoE Optimization for Semantic Self-Correcting Video Transmission in Multi-UAV Networks
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
TVMC: Time-Varying Mesh Compression via Multi-Stage Anchor Mesh Generation
by: Huang, He, et al.
Published: (2025)
by: Huang, He, et al.
Published: (2025)
ABC: Adaptive BayesNet Structure Learning for Computational Scalable Multi-task Image Compression
by: Zhang, Yufeng, et al.
Published: (2025)
by: Zhang, Yufeng, et al.
Published: (2025)
A H.265/HEVC Fine-Grained ROI Video Encryption Algorithm Based on Coding Unit and Prompt Segmentation
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
by: Shimada, Kazuki, et al.
Published: (2025)
by: Shimada, Kazuki, et al.
Published: (2025)
Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
by: Sulun, Serkan, et al.
Published: (2025)
by: Sulun, Serkan, et al.
Published: (2025)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
by: Wen, Liuyuan
Published: (2024)
by: Wen, Liuyuan
Published: (2024)
NiMark: A Non-intrusive Watermarking Framework against Screen-shooting Attacks
by: Wu, Yufeng, et al.
Published: (2026)
by: Wu, Yufeng, et al.
Published: (2026)
Generative Flow Networks for Personalized Multimedia Systems: A Case Study on Short Video Feeds
by: Jin, Yili, et al.
Published: (2025)
by: Jin, Yili, et al.
Published: (2025)
Enhanced Template-based Intra Mode Derivation with Adaptive Block Vector Replacement
by: Zhang, Jiaqi, et al.
Published: (2025)
by: Zhang, Jiaqi, et al.
Published: (2025)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
Smaller is Better: Generative Models Can Power Short Video Preloading
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Dynamic resolution switching for live streaming
by: Xiong, Xin, et al.
Published: (2026)
by: Xiong, Xin, et al.
Published: (2026)
EyeNexus: Adaptive Gaze-Driven Quality and Bitrate Streaming for Seamless VR Cloud Gaming Experiences
by: Wu, Ze, et al.
Published: (2025)
by: Wu, Ze, et al.
Published: (2025)
HybridPrompt: Bridging Generative Priors and Traditional Codecs for Mobile Streaming
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Learning Perceptual Representations for Gaming NR-VQA with Multi-Task FR Signals
by: Chen, Yu-Chih, et al.
Published: (2026)
by: Chen, Yu-Chih, et al.
Published: (2026)
Dual Inverse Degradation Network for Real-World SDRTV-to-HDRTV Conversion
by: Xu, Kepeng, et al.
Published: (2023)
by: Xu, Kepeng, et al.
Published: (2023)
A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
by: Wu, Chengyu, et al.
Published: (2024)
by: Wu, Chengyu, et al.
Published: (2024)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Enhanced Quality Aware-Scalable Underwater Image Compression
by: Zhu, Linwei, et al.
Published: (2025)
by: Zhu, Linwei, et al.
Published: (2025)
Avoiding Quality Saturation in UGC Compression Using Denoised References
by: Xiong, Xin, et al.
Published: (2025)
by: Xiong, Xin, et al.
Published: (2025)
Similar Items
-
Computation-resource-efficient Task-oriented Communications
by: Fu, Jingwen, et al.
Published: (2025) -
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023) -
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024) -
DriftDecode: One-Step Wireless Image Decoding via Drifting-Inspired Detail Recovery
by: Fu, Jingwen, et al.
Published: (2026) -
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
by: Ragano, Alessandro, et al.
Published: (2024)