VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Torgashov, Nikita, Henter, Gustav Eje, Skantze, Gabriel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908787739197440
author Torgashov, Nikita
Henter, Gustav Eje
Skantze, Gabriel
author_facet Torgashov, Nikita
Henter, Gustav Eje
Skantze, Gabriel
contents We present VoXtream, a fully autoregressive, zero-shot streaming text-to-speech (TTS) system for real-time use that begins speaking from the first word. VoXtream directly maps incoming phonemes to audio tokens using a monotonic alignment scheme and a limited look-ahead that does not delay onset. Built around an incremental phoneme transformer, a temporal transformer predicting semantic and duration tokens, and a depth transformer producing acoustic tokens, VoXtream achieves, to our knowledge, the lowest initial delay among publicly available streaming TTS: 102 ms on GPU. Despite being trained on a mid-scale 9k-hour corpus, it matches or surpasses larger baselines on several metrics, while delivering competitive quality in both output- and full-streaming settings. Demo and code are available at https://herimor.github.io/voxtream.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15969
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
Torgashov, Nikita
Henter, Gustav Eje
Skantze, Gabriel
Audio and Speech Processing
Computation and Language
Human-Computer Interaction
Machine Learning
Sound
We present VoXtream, a fully autoregressive, zero-shot streaming text-to-speech (TTS) system for real-time use that begins speaking from the first word. VoXtream directly maps incoming phonemes to audio tokens using a monotonic alignment scheme and a limited look-ahead that does not delay onset. Built around an incremental phoneme transformer, a temporal transformer predicting semantic and duration tokens, and a depth transformer producing acoustic tokens, VoXtream achieves, to our knowledge, the lowest initial delay among publicly available streaming TTS: 102 ms on GPU. Despite being trained on a mid-scale 9k-hour corpus, it matches or surpasses larger baselines on several metrics, while delivering competitive quality in both output- and full-streaming settings. Demo and code are available at https://herimor.github.io/voxtream.
title VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
topic Audio and Speech Processing
Computation and Language
Human-Computer Interaction
Machine Learning
Sound
url https://arxiv.org/abs/2509.15969