Error-Resilient Semantic Communication for Speech Transmission over Packet-Loss Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Zhuohang, Dai, Jincheng, Yao, Shengshi, Wang, Junyi, Li, Yanlong, Niu, Kai, Xu, Wenjun, Zhang, Ping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
Semantic MIMO Systems for Speech-to-Text Transmission
di: Weng, Zhenzi, et al.
Pubblicazione: (2024)
di: Weng, Zhenzi, et al.
Pubblicazione: (2024)
Large Speech Model Enabled Semantic Communication
di: Tian, Yun, et al.
Pubblicazione: (2025)
di: Tian, Yun, et al.
Pubblicazione: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning
di: Wang, Junyi, et al.
Pubblicazione: (2026)
di: Wang, Junyi, et al.
Pubblicazione: (2026)
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
The IEEE-IS2 2024 Music Packet Loss Concealment Challenge
di: Mezza, Alessandro Ilic, et al.
Pubblicazione: (2024)
di: Mezza, Alessandro Ilic, et al.
Pubblicazione: (2024)
The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge
di: Diener, Lorenz, et al.
Pubblicazione: (2024)
di: Diener, Lorenz, et al.
Pubblicazione: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
Semantic Communications for Speech Recognition
di: Weng, Zhenzi, et al.
Pubblicazione: (2021)
di: Weng, Zhenzi, et al.
Pubblicazione: (2021)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning
di: Wang, Junyi, et al.
Pubblicazione: (2026)
di: Wang, Junyi, et al.
Pubblicazione: (2026)
Generative Semantic Communication for Text-to-Speech Synthesis
di: Zheng, Jiahao, et al.
Pubblicazione: (2024)
di: Zheng, Jiahao, et al.
Pubblicazione: (2024)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
di: Cheng, Changhao, et al.
Pubblicazione: (2026)
di: Cheng, Changhao, et al.
Pubblicazione: (2026)
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
di: Dong, Zhongren, et al.
Pubblicazione: (2025)
di: Dong, Zhongren, et al.
Pubblicazione: (2025)
From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation
di: Xue, Ke, et al.
Pubblicazione: (2025)
di: Xue, Ke, et al.
Pubblicazione: (2025)
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
di: Liu, Yanyan, et al.
Pubblicazione: (2025)
di: Liu, Yanyan, et al.
Pubblicazione: (2025)
RADE for Land Mobile Radio: A Neural Codec for Transmission of Speech over Baseband FM Radio Channels
di: Rowe, David, et al.
Pubblicazione: (2025)
di: Rowe, David, et al.
Pubblicazione: (2025)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
RTCFake: Speech Deepfake Detection in Real-Time Communication
di: Xue, Jun, et al.
Pubblicazione: (2026)
di: Xue, Jun, et al.
Pubblicazione: (2026)
Deep Generative Modeling Reshapes Compression and Transmission: From Efficiency to Resiliency
di: Dai, Jincheng, et al.
Pubblicazione: (2024)
di: Dai, Jincheng, et al.
Pubblicazione: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
di: Guo, Hao-Han, et al.
Pubblicazione: (2024)
di: Guo, Hao-Han, et al.
Pubblicazione: (2024)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
di: Liu, Qingyu, et al.
Pubblicazione: (2025)
di: Liu, Qingyu, et al.
Pubblicazione: (2025)
Speaker Recognition -- Wavelet Packet Based Multiresolution Feature Extraction Approach
di: Bhardwaj, Saurabh, et al.
Pubblicazione: (2025)
di: Bhardwaj, Saurabh, et al.
Pubblicazione: (2025)
Loud-loss: A Perceptually Motivated Loss Function for Speech Enhancement Based on Equal-Loudness Contours
di: Li, Zixuan, et al.
Pubblicazione: (2025)
di: Li, Zixuan, et al.
Pubblicazione: (2025)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
A MATLAB toolbox for Computation of Speech Transmission Index (STI)
di: Rajmic, Pavel, et al.
Pubblicazione: (2025)
di: Rajmic, Pavel, et al.
Pubblicazione: (2025)
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
di: Li, Longhao, et al.
Pubblicazione: (2025)
di: Li, Longhao, et al.
Pubblicazione: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
Dynamic Fusion Multimodal Network for SpeechWellness Detection
di: Sun, Wenqiang, et al.
Pubblicazione: (2025)
di: Sun, Wenqiang, et al.
Pubblicazione: (2025)
HQ-MPSD: A Multilingual Artifact-Controlled Benchmark for Partial Deepfake Speech Detection
di: Li, Menglu, et al.
Pubblicazione: (2025)
di: Li, Menglu, et al.
Pubblicazione: (2025)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
di: Du, Yichao, et al.
Pubblicazione: (2024)
di: Du, Yichao, et al.
Pubblicazione: (2024)
Physics-Informed Neural Networks for Speech Production
di: Yokota, Kazuya, et al.
Pubblicazione: (2025)
di: Yokota, Kazuya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
di: Yao, Shengshi, et al.
Pubblicazione: (2025) -
Semantic MIMO Systems for Speech-to-Text Transmission
di: Weng, Zhenzi, et al.
Pubblicazione: (2024) -
Large Speech Model Enabled Semantic Communication
di: Tian, Yun, et al.
Pubblicazione: (2025) -
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025) -
Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning
di: Wang, Junyi, et al.
Pubblicazione: (2026)