A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xinyi, Huang, Danlan, Qi, Zhixin, Zhang, Liang, Jiang, Ting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909391063613440
author Zhou, Xinyi
Huang, Danlan
Qi, Zhixin
Zhang, Liang
Jiang, Ting
author_facet Zhou, Xinyi
Huang, Danlan
Qi, Zhixin
Zhang, Liang
Jiang, Ting
contents Deep joint source-channel coding (DeepJSCC) has shown promise in wireless transmission of text, speech, and images within the realm of semantic communication. However, wireless video transmission presents greater challenges due to the difficulty of extracting and compactly representing both spatial and temporal features, as well as its significant bandwidth and computational resource requirements. In response, we propose a novel video DeepJSCC (VDJSCC) approach to enable end-to-end video transmission over a wireless channel. Our approach involves the design of a multi-scale vision Transformer encoder and decoder to effectively capture spatial-temporal representations over long-term frames. Additionally, we propose a dynamic token selection module to mask less semantically important tokens from spatial or temporal dimensions, allowing for content-adaptive variable-length video coding by adjusting the token keep ratio. Experimental results demonstrate the effectiveness of our VDJSCC approach compared to digital schemes that use separate source and channel codes, as well as other DeepJSCC schemes, in terms of reconstruction quality and bandwidth reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09936
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission
Zhou, Xinyi
Huang, Danlan
Qi, Zhixin
Zhang, Liang
Jiang, Ting
Image and Video Processing
Deep joint source-channel coding (DeepJSCC) has shown promise in wireless transmission of text, speech, and images within the realm of semantic communication. However, wireless video transmission presents greater challenges due to the difficulty of extracting and compactly representing both spatial and temporal features, as well as its significant bandwidth and computational resource requirements. In response, we propose a novel video DeepJSCC (VDJSCC) approach to enable end-to-end video transmission over a wireless channel. Our approach involves the design of a multi-scale vision Transformer encoder and decoder to effectively capture spatial-temporal representations over long-term frames. Additionally, we propose a dynamic token selection module to mask less semantically important tokens from spatial or temporal dimensions, allowing for content-adaptive variable-length video coding by adjusting the token keep ratio. Experimental results demonstrate the effectiveness of our VDJSCC approach compared to digital schemes that use separate source and channel codes, as well as other DeepJSCC schemes, in terms of reconstruction quality and bandwidth reduction.
title A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission
topic Image and Video Processing
url https://arxiv.org/abs/2411.09936