Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Feibo, Tu, Siwei, Dong, Li, Pan, Cunhua, Wang, Jiangzhou, You, Xiaohu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2411.03876
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909379086778368
author Jiang, Feibo
Tu, Siwei
Dong, Li
Pan, Cunhua
Wang, Jiangzhou
You, Xiaohu
author_facet Jiang, Feibo
Tu, Siwei
Dong, Li
Pan, Cunhua
Wang, Jiangzhou
You, Xiaohu
contents The rapid development of generative Artificial Intelligence (AI) continually unveils the potential of Semantic Communication (SemCom). However, current talking-face SemCom systems still encounter challenges such as low bandwidth utilization, semantic ambiguity, and diminished Quality of Experience (QoE). This study introduces a Large Generative Model-assisted Talking-face Semantic Communication (LGM-TSC) System tailored for the talking-face video communication. Firstly, we introduce a Generative Semantic Extractor (GSE) at the transmitter based on the FunASR model to convert semantically sparse talking-face videos into texts with high information density. Secondly, we establish a private Knowledge Base (KB) based on the Large Language Model (LLM) for semantic disambiguation and correction, complemented by a joint knowledge base-semantic-channel coding scheme. Finally, at the receiver, we propose a Generative Semantic Reconstructor (GSR) that utilizes BERT-VITS2 and SadTalker models to transform text back into a high-QoE talking-face video matching the user's timbre. Simulation results demonstrate the feasibility and effectiveness of the proposed LGM-TSC system.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03876
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Generative Model-assisted Talking-face Semantic Communication System
Jiang, Feibo
Tu, Siwei
Dong, Li
Pan, Cunhua
Wang, Jiangzhou
You, Xiaohu
Information Theory
Machine Learning
The rapid development of generative Artificial Intelligence (AI) continually unveils the potential of Semantic Communication (SemCom). However, current talking-face SemCom systems still encounter challenges such as low bandwidth utilization, semantic ambiguity, and diminished Quality of Experience (QoE). This study introduces a Large Generative Model-assisted Talking-face Semantic Communication (LGM-TSC) System tailored for the talking-face video communication. Firstly, we introduce a Generative Semantic Extractor (GSE) at the transmitter based on the FunASR model to convert semantically sparse talking-face videos into texts with high information density. Secondly, we establish a private Knowledge Base (KB) based on the Large Language Model (LLM) for semantic disambiguation and correction, complemented by a joint knowledge base-semantic-channel coding scheme. Finally, at the receiver, we propose a Generative Semantic Reconstructor (GSR) that utilizes BERT-VITS2 and SadTalker models to transform text back into a high-QoE talking-face video matching the user's timbre. Simulation results demonstrate the feasibility and effectiveness of the proposed LGM-TSC system.
title Large Generative Model-assisted Talking-face Semantic Communication System
topic Information Theory
Machine Learning
url https://arxiv.org/abs/2411.03876