SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Shengshi, Dai, Jincheng, Qin, Xiaoqi, Wang, Sixian, Wang, Siye, Niu, Kai, Zhang, Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929684001849344
author Yao, Shengshi
Dai, Jincheng
Qin, Xiaoqi
Wang, Sixian
Wang, Siye
Niu, Kai
Zhang, Ping
author_facet Yao, Shengshi
Dai, Jincheng
Qin, Xiaoqi
Wang, Sixian
Wang, Siye
Niu, Kai
Zhang, Ping
contents In this paper, we propose "SoundSpring", a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems. Unlike recent deep JSCC transceivers, which learn to directly map audio signals to analog channel-input symbols via neural networks, our SoundSpring adopts the layered architecture that delineates audio compression from digital coded transmission, but it sufficiently exploits the impressive in-context predictive capabilities of large language (foundation) models. Integrated with the casual-order mask learning strategy, our single model operates on the latent feature domain and serve dual-functionalities: as efficient audio compressors at the transmitter and as effective mechanisms for packet loss concealment at the receiver. By jointly optimizing towards both audio compression efficiency and transmission error resiliency, we show that mask-learned language models are indeed powerful contextual predictors, and our dual-functional compression and concealment framework offers fresh perspectives on the application of foundation language models in audio communication. Through extensive experimental evaluations, we establish that SoundSpring apparently outperforms contemporary audio transmission systems in terms of signal fidelity metrics and perceptual quality scores. These new findings not only advocate for the practical deployment of SoundSpring in learning-based audio communication systems but also inspire the development of future audio semantic transceivers.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
Yao, Shengshi
Dai, Jincheng
Qin, Xiaoqi
Wang, Sixian
Wang, Siye
Niu, Kai
Zhang, Ping
Audio and Speech Processing
Sound
Signal Processing
In this paper, we propose "SoundSpring", a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems. Unlike recent deep JSCC transceivers, which learn to directly map audio signals to analog channel-input symbols via neural networks, our SoundSpring adopts the layered architecture that delineates audio compression from digital coded transmission, but it sufficiently exploits the impressive in-context predictive capabilities of large language (foundation) models. Integrated with the casual-order mask learning strategy, our single model operates on the latent feature domain and serve dual-functionalities: as efficient audio compressors at the transmitter and as effective mechanisms for packet loss concealment at the receiver. By jointly optimizing towards both audio compression efficiency and transmission error resiliency, we show that mask-learned language models are indeed powerful contextual predictors, and our dual-functional compression and concealment framework offers fresh perspectives on the application of foundation language models in audio communication. Through extensive experimental evaluations, we establish that SoundSpring apparently outperforms contemporary audio transmission systems in terms of signal fidelity metrics and perceptual quality scores. These new findings not only advocate for the practical deployment of SoundSpring in learning-based audio communication systems but also inspire the development of future audio semantic transceivers.
title SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2501.12696