StyleStream: Real-Time Zero-Shot Voice Style Conversion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yisi, Lee, Nicholas, Anumanchipalli, Gopala
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915812397285376
author Liu, Yisi
Lee, Nicholas
Anumanchipalli, Gopala
author_facet Liu, Yisi
Lee, Nicholas
Anumanchipalli, Gopala
contents Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work has explored this problem, conversion quality remains limited, and real-time voice style conversion has not been addressed. We propose StyleStream, the first streamable zero-shot voice style conversion system that achieves state-of-the-art performance. StyleStream consists of two components: a Destylizer, which removes style attributes while preserving linguistic content, and a Stylizer, a diffusion transformer (DiT) that reintroduces target style conditioned on reference speech. Robust content-style disentanglement is enforced through text supervision and a highly constrained information bottleneck. This design enables a fully non-autoregressive architecture, achieving real-time voice style conversion with an end-to-end latency of 1 second. Samples and real-time demo: https://berkeley-speech-group.github.io/StyleStream/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StyleStream: Real-Time Zero-Shot Voice Style Conversion
Liu, Yisi
Lee, Nicholas
Anumanchipalli, Gopala
Sound
Artificial Intelligence
Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work has explored this problem, conversion quality remains limited, and real-time voice style conversion has not been addressed. We propose StyleStream, the first streamable zero-shot voice style conversion system that achieves state-of-the-art performance. StyleStream consists of two components: a Destylizer, which removes style attributes while preserving linguistic content, and a Stylizer, a diffusion transformer (DiT) that reintroduces target style conditioned on reference speech. Robust content-style disentanglement is enforced through text supervision and a highly constrained information bottleneck. This design enables a fully non-autoregressive architecture, achieving real-time voice style conversion with an end-to-end latency of 1 second. Samples and real-time demo: https://berkeley-speech-group.github.io/StyleStream/.
title StyleStream: Real-Time Zero-Shot Voice Style Conversion
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2602.20113