Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Junchuan, Wang, Xintong, Wang, Ye
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916973916454912
author Zhao, Junchuan
Wang, Xintong
Wang, Ye
author_facet Zhao, Junchuan
Wang, Xintong
Wang, Ye
contents Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice conversion (VC) model within the VALLE-X framework, leveraging its strong in-context learning capabilities for speaker adaptation. To enhance prosody control, we introduce a prosody-aware audio codec encoder (PACE) module, which isolates and refines prosody from other sources, improving expressiveness and control. By integrating PACE into our VC model, we achieve greater flexibility in prosody manipulation while preserving speaker timbre. Experimental evaluation results demonstrate that our approach outperforms baseline VC systems in prosody preservation, timbre consistency, and overall naturalness, surpassing baseline VC systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
Zhao, Junchuan
Wang, Xintong
Wang, Ye
Sound
Artificial Intelligence
Audio and Speech Processing
Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice conversion (VC) model within the VALLE-X framework, leveraging its strong in-context learning capabilities for speaker adaptation. To enhance prosody control, we introduce a prosody-aware audio codec encoder (PACE) module, which isolates and refines prosody from other sources, improving expressiveness and control. By integrating PACE into our VC model, we achieve greater flexibility in prosody manipulation while preserving speaker timbre. Experimental evaluation results demonstrate that our approach outperforms baseline VC systems in prosody preservation, timbre consistency, and overall naturalness, surpassing baseline VC systems.
title Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.15402