Large Model Empowered Streaming Speech Semantic Communications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Weng, Zhenzi, Qin, Zhijin, Li, Geoffrey Ye
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916624795172864
author Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
author_facet Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
contents In this paper, we introduce a large model-empowered streaming semantic communication system for speech transmission across various languages, named LSSC-ST. Specifically, we devise an edge-device collaborative semantic communication architecture by offloading the intricate semantic extraction and channel coding modules to edge servers, thereby reducing the computational burden on local devices. To support multilingual speech transmission, pre-trained large speech models are utilized to learn unified semantic features from speech in different languages, breaking the constraint of a single input language and enhancing the practicality of the LSSC-ST. Moreover, the input speech is sequentially streamed into the developed system as short speech segments, which enables low transmission latency without degrading the quality of the produced speech. A novel dynamic speech segmentation algorithm is proposed to further reduce the transmission latency by adaptively adjusting the duration of speech segments. According to simulation results, the LSSC-ST provides more accurate speech transmission and achieves a streaming manner with lower latency compared to the existing non-streaming semantic communication systems.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Model Empowered Streaming Speech Semantic Communications
Weng, Zhenzi
Qin, Zhijin
Li, Geoffrey Ye
Audio and Speech Processing
In this paper, we introduce a large model-empowered streaming semantic communication system for speech transmission across various languages, named LSSC-ST. Specifically, we devise an edge-device collaborative semantic communication architecture by offloading the intricate semantic extraction and channel coding modules to edge servers, thereby reducing the computational burden on local devices. To support multilingual speech transmission, pre-trained large speech models are utilized to learn unified semantic features from speech in different languages, breaking the constraint of a single input language and enhancing the practicality of the LSSC-ST. Moreover, the input speech is sequentially streamed into the developed system as short speech segments, which enables low transmission latency without degrading the quality of the produced speech. A novel dynamic speech segmentation algorithm is proposed to further reduce the transmission latency by adaptively adjusting the duration of speech segments. According to simulation results, the LSSC-ST provides more accurate speech transmission and achieves a streaming manner with lower latency compared to the existing non-streaming semantic communication systems.
title Large Model Empowered Streaming Speech Semantic Communications
topic Audio and Speech Processing
url https://arxiv.org/abs/2501.05859