Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Bin, Yu, Wenbo, Zhang, Qinshan, Zhuang, Tianqu, Wu, Hao, Jiang, Yong, Xia, Shu-Tao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910077141647360
author Chen, Bin
Yu, Wenbo
Zhang, Qinshan
Zhuang, Tianqu
Wu, Hao
Jiang, Yong
Xia, Shu-Tao
author_facet Chen, Bin
Yu, Wenbo
Zhang, Qinshan
Zhuang, Tianqu
Wu, Hao
Jiang, Yong
Xia, Shu-Tao
contents Interactive computer vision (CV) plays a crucial role in various real-world applications, whose performance is highly dependent on communication networks. Nonetheless, the data-oriented characteristics of conventional communications often do not align with the special needs of interactive CV tasks. To alleviate this issue, the recently emerged semantic communications only transmit task-related semantic information and exhibit a promising landscape to address this problem. However, the communication challenges associated with Semantic Facial Editing, one of the most important interactive CV applications on social media, still remain largely unexplored. In this paper, we fill this gap by proposing Editable-DeepSC, a novel cross-modal semantic communication approach for facial editing. Firstly, we theoretically discuss different transmission schemes that separately handle communications and editings, and emphasize the necessity of Joint Editing-Channel Coding (JECC) via iterative attributes matching, which integrates editings into the communication chain to preserve more semantic mutual information. To compactly represent the high-dimensional data, we leverage inversion methods via pre-trained StyleGAN priors for semantic coding. To tackle the dynamic channel noise conditions, we propose SNR-aware channel coding via model fine-tuning. Extensive experiments indicate that Editable-DeepSC can achieve superior editings while significantly saving the transmission bandwidth, even under high-resolution and out-of-distribution (OOD) settings.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15702
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
Chen, Bin
Yu, Wenbo
Zhang, Qinshan
Zhuang, Tianqu
Wu, Hao
Jiang, Yong
Xia, Shu-Tao
Information Theory
Computer Vision and Pattern Recognition
Networking and Internet Architecture
Interactive computer vision (CV) plays a crucial role in various real-world applications, whose performance is highly dependent on communication networks. Nonetheless, the data-oriented characteristics of conventional communications often do not align with the special needs of interactive CV tasks. To alleviate this issue, the recently emerged semantic communications only transmit task-related semantic information and exhibit a promising landscape to address this problem. However, the communication challenges associated with Semantic Facial Editing, one of the most important interactive CV applications on social media, still remain largely unexplored. In this paper, we fill this gap by proposing Editable-DeepSC, a novel cross-modal semantic communication approach for facial editing. Firstly, we theoretically discuss different transmission schemes that separately handle communications and editings, and emphasize the necessity of Joint Editing-Channel Coding (JECC) via iterative attributes matching, which integrates editings into the communication chain to preserve more semantic mutual information. To compactly represent the high-dimensional data, we leverage inversion methods via pre-trained StyleGAN priors for semantic coding. To tackle the dynamic channel noise conditions, we propose SNR-aware channel coding via model fine-tuning. Extensive experiments indicate that Editable-DeepSC can achieve superior editings while significantly saving the transmission bandwidth, even under high-resolution and out-of-distribution (OOD) settings.
title Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
topic Information Theory
Computer Vision and Pattern Recognition
Networking and Internet Architecture
url https://arxiv.org/abs/2411.15702