Robust Semantic Communications for Speech Transmission

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Weng, Zhenzi, Qin, Zhijin, Li, Geoffrey Ye
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911037292281856
author Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
author_facet Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
contents In this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Specifically, we consider the speech-to-text translation (S2TT) as the transmission goal. First, a new deep semantic encoder is developed to convert speech in the source language to textual features associated with the target language, facilitating the end-to-end semantic exchange to perform the S2TT task and reducing the transmission data without performance degradation. To mitigate semantic impairments inherent in the corrupted speech, a novel generative adversarial network (GAN)-enabled deep semantic compensator is established to estimate the lost semantic information within the speech and extract deep semantic features simultaneously, which enables robust semantic transmission for corrupted speech. Furthermore, a semantic probe-aided compensator is devised to enhance the semantic fidelity of recovered semantic features and improve the understandability of the target text. According to simulation results, the proposed Ross-S2T exhibits superior S2TT performance compared to conventional approaches and high robustness against semantic impairments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Semantic Communications for Speech Transmission
Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
Audio and Speech Processing
In this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Specifically, we consider the speech-to-text translation (S2TT) as the transmission goal. First, a new deep semantic encoder is developed to convert speech in the source language to textual features associated with the target language, facilitating the end-to-end semantic exchange to perform the S2TT task and reducing the transmission data without performance degradation. To mitigate semantic impairments inherent in the corrupted speech, a novel generative adversarial network (GAN)-enabled deep semantic compensator is established to estimate the lost semantic information within the speech and extract deep semantic features simultaneously, which enables robust semantic transmission for corrupted speech. Furthermore, a semantic probe-aided compensator is devised to enhance the semantic fidelity of recovered semantic features and improve the understandability of the target text. According to simulation results, the proposed Ross-S2T exhibits superior S2TT performance compared to conventional approaches and high robustness against semantic impairments.
title Robust Semantic Communications for Speech Transmission
topic Audio and Speech Processing
url https://arxiv.org/abs/2403.05187