Simultaneous Speech-to-Speech Translation Without Aligned Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Labiausse, Tom, Fabre, Romain, Estève, Yannick, Défossez, Alexandre, Zeghidour, Neil
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914323299827712
author Labiausse, Tom
Fabre, Romain
Estève, Yannick
Défossez, Alexandre
Zeghidour, Neil
author_facet Labiausse, Tom
Fabre, Romain
Estève, Yannick
Défossez, Alexandre
Zeghidour, Neil
contents Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencies. Traditional approaches rely on supervised training with word-level aligned data, which is difficult to collect at scale and thus depends on synthetic alignments using language-specific heuristics that are suboptimal. We propose Hibiki-Zero, which eliminates the need for word-level alignments entirely. This fundamentally simplifies the training pipeline and enables seamless scaling to diverse languages with varying grammatical structures, removing the bottleneck of designing language-specific alignment heuristics. We first train on sentence-level aligned data to learn speech translation at high latency, then apply a novel reinforcement learning strategy using GRPO to optimize latency while preserving translation quality. Hibiki-Zero achieves state-of-the-art performance in translation accuracy, latency, voice transfer, and naturalness across five X-to-English tasks. Moreover, we demonstrate that our model can be adapted to support a new input language with less than 1000h of speech. We provide examples, model weights, inference code and we release a benchmark containing 45h of multilingual data for speech translation evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11072
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Simultaneous Speech-to-Speech Translation Without Aligned Data
Labiausse, Tom
Fabre, Romain
Estève, Yannick
Défossez, Alexandre
Zeghidour, Neil
Computation and Language
Sound
Audio and Speech Processing
Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencies. Traditional approaches rely on supervised training with word-level aligned data, which is difficult to collect at scale and thus depends on synthetic alignments using language-specific heuristics that are suboptimal. We propose Hibiki-Zero, which eliminates the need for word-level alignments entirely. This fundamentally simplifies the training pipeline and enables seamless scaling to diverse languages with varying grammatical structures, removing the bottleneck of designing language-specific alignment heuristics. We first train on sentence-level aligned data to learn speech translation at high latency, then apply a novel reinforcement learning strategy using GRPO to optimize latency while preserving translation quality. Hibiki-Zero achieves state-of-the-art performance in translation accuracy, latency, voice transfer, and naturalness across five X-to-English tasks. Moreover, we demonstrate that our model can be adapted to support a new input language with less than 1000h of speech. We provide examples, model weights, inference code and we release a benchmark containing 45h of multilingual data for speech translation evaluation.
title Simultaneous Speech-to-Speech Translation Without Aligned Data
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2602.11072