Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915297643986944 |
|---|---|
| author | Lu, Ye-Xin Du, Hui-Peng Liu, Fei Ai, Yang Ling, Zhen-Hua |
| author_facet | Lu, Ye-Xin Du, Hui-Peng Liu, Fei Ai, Yang Ling, Zhen-Hua |
| contents | Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13830 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising Lu, Ye-Xin Du, Hui-Peng Liu, Fei Ai, Yang Ling, Zhen-Hua Audio and Speech Processing Sound Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models. |
| title | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.13830 |