Rate-Distortion-Perception Controllable Joint Source-Channel Coding for High-Fidelity Generative Communications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Kailin, Dai, Jincheng, Liu, Zhenyu, Wang, Sixian, Qin, Xiaoqi, Xu, Wenjun, Niu, Kai, Zhang, Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912002175139840
author Tan, Kailin
Dai, Jincheng
Liu, Zhenyu
Wang, Sixian
Qin, Xiaoqi
Xu, Wenjun
Niu, Kai
Zhang, Ping
author_facet Tan, Kailin
Dai, Jincheng
Liu, Zhenyu
Wang, Sixian
Qin, Xiaoqi
Xu, Wenjun
Niu, Kai
Zhang, Ping
contents End-to-end image transmission has recently become a crucial trend in intelligent wireless communications, driven by the increasing demand for high bandwidth efficiency. However, existing methods primarily optimize the trade-off between bandwidth cost and objective distortion, often failing to deliver visually pleasing results aligned with human perception. In this paper, we propose a novel rate-distortion-perception (RDP) jointly optimized joint source-channel coding (JSCC) framework to enhance perception quality in human communications. Our RDP-JSCC framework integrates a flexible plug-in conditional Generative Adversarial Networks (GANs) to provide detailed and realistic image reconstructions at the receiver, overcoming the limitations of traditional rate-distortion optimized solutions that typically produce blurry or poorly textured images. Based on this framework, we introduce a distortion-perception controllable transmission (DPCT) model, which addresses the variation in the perception-distortion trade-off. DPCT uses a lightweight spatial realism embedding module (SREM) to condition the generator on a realism map, enabling the customization of appearance realism for each image region at the receiver from a single transmission. Furthermore, for scenarios with scarce bandwidth, we propose an interest-oriented content-controllable transmission (CCT) model. CCT prioritizes the transmission of regions that attract user attention and generates other regions from an instance label map, ensuring both content consistency and appearance realism for all regions while proportionally reducing channel bandwidth costs. Comprehensive experiments demonstrate the superiority of our RDP-optimized image transmission framework over state-of-the-art engineered image transmission systems and advanced perceptual methods.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rate-Distortion-Perception Controllable Joint Source-Channel Coding for High-Fidelity Generative Communications
Tan, Kailin
Dai, Jincheng
Liu, Zhenyu
Wang, Sixian
Qin, Xiaoqi
Xu, Wenjun
Niu, Kai
Zhang, Ping
Image and Video Processing
End-to-end image transmission has recently become a crucial trend in intelligent wireless communications, driven by the increasing demand for high bandwidth efficiency. However, existing methods primarily optimize the trade-off between bandwidth cost and objective distortion, often failing to deliver visually pleasing results aligned with human perception. In this paper, we propose a novel rate-distortion-perception (RDP) jointly optimized joint source-channel coding (JSCC) framework to enhance perception quality in human communications. Our RDP-JSCC framework integrates a flexible plug-in conditional Generative Adversarial Networks (GANs) to provide detailed and realistic image reconstructions at the receiver, overcoming the limitations of traditional rate-distortion optimized solutions that typically produce blurry or poorly textured images. Based on this framework, we introduce a distortion-perception controllable transmission (DPCT) model, which addresses the variation in the perception-distortion trade-off. DPCT uses a lightweight spatial realism embedding module (SREM) to condition the generator on a realism map, enabling the customization of appearance realism for each image region at the receiver from a single transmission. Furthermore, for scenarios with scarce bandwidth, we propose an interest-oriented content-controllable transmission (CCT) model. CCT prioritizes the transmission of regions that attract user attention and generates other regions from an instance label map, ensuring both content consistency and appearance realism for all regions while proportionally reducing channel bandwidth costs. Comprehensive experiments demonstrate the superiority of our RDP-optimized image transmission framework over state-of-the-art engineered image transmission systems and advanced perceptual methods.
title Rate-Distortion-Perception Controllable Joint Source-Channel Coding for High-Fidelity Generative Communications
topic Image and Video Processing
url https://arxiv.org/abs/2408.14127