Latent Speech-Text Transformer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lu, Yen-Ju, Gaur, Yashesh, Zhou, Wei, Muller, Benjamin, Villalba, Jesus, Dehak, Najim, Zettlemoyer, Luke, Ghosh, Gargi, Lewis, Mike, Iyer, Srinivasan, Le, Duc
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912956961259520
author Lu, Yen-Ju
Gaur, Yashesh
Zhou, Wei
Muller, Benjamin
Villalba, Jesus
Dehak, Najim
Zettlemoyer, Luke
Ghosh, Gargi
Lewis, Mike
Iyer, Srinivasan
Le, Duc
author_facet Lu, Yen-Ju
Gaur, Yashesh
Zhou, Wei
Muller, Benjamin
Villalba, Jesus
Dehak, Najim
Zettlemoyer, Luke
Ghosh, Gargi
Lewis, Mike
Iyer, Srinivasan
Le, Duc
contents Auto-regressive speech-text models pre-trained on interleaved text tokens and discretized speech tokens demonstrate strong speech understanding and generation, yet remain substantially less compute-efficient than text LLMs, partly due to the much longer sequences of speech tokens relative to text. This modality imbalance disproportionately allocates pre-training and inference compute to speech, potentially hindering effective cross-modal alignment and slowing performance scaling by orders of magnitude. We introduce the Latent Speech-Text Transformer (LST), which aggregates speech tokens into latent speech patches that serve as higher-level autoregressive units. This design aligns the sequence-modeling granularity between speech and text while improving computational efficiency. The resulting patches can align with textual units to facilitate cross-modal knowledge transfer and compactly capture recurring acoustic patterns such as silence. Across story-completion benchmarks under both compute-controlled and data-controlled settings, LST consistently improves speech accuracy while also improving text performance, achieving up to +6.5% absolute gain on speech HellaSwag in compute-controlled training (+5.3% in data-controlled training). Under compute-controlled scaling from 420M to 1.8B parameters in a near compute-optimal regime, gains grow with scale, and improvements persist up to 7B parameters under fixed-token budgets. These benefits extend to downstream tasks: LST stabilizes ASR adaptation and reduces the effective autoregressive sequence length during ASR and TTS inference, lowering computational cost without degrading reconstruction quality. The code is available at https://github.com/facebookresearch/lst.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Speech-Text Transformer
Lu, Yen-Ju
Gaur, Yashesh
Zhou, Wei
Muller, Benjamin
Villalba, Jesus
Dehak, Najim
Zettlemoyer, Luke
Ghosh, Gargi
Lewis, Mike
Iyer, Srinivasan
Le, Duc
Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Auto-regressive speech-text models pre-trained on interleaved text tokens and discretized speech tokens demonstrate strong speech understanding and generation, yet remain substantially less compute-efficient than text LLMs, partly due to the much longer sequences of speech tokens relative to text. This modality imbalance disproportionately allocates pre-training and inference compute to speech, potentially hindering effective cross-modal alignment and slowing performance scaling by orders of magnitude. We introduce the Latent Speech-Text Transformer (LST), which aggregates speech tokens into latent speech patches that serve as higher-level autoregressive units. This design aligns the sequence-modeling granularity between speech and text while improving computational efficiency. The resulting patches can align with textual units to facilitate cross-modal knowledge transfer and compactly capture recurring acoustic patterns such as silence. Across story-completion benchmarks under both compute-controlled and data-controlled settings, LST consistently improves speech accuracy while also improving text performance, achieving up to +6.5% absolute gain on speech HellaSwag in compute-controlled training (+5.3% in data-controlled training). Under compute-controlled scaling from 420M to 1.8B parameters in a near compute-optimal regime, gains grow with scale, and improvements persist up to 7B parameters under fixed-token budgets. These benefits extend to downstream tasks: LST stabilizes ASR adaptation and reduces the effective autoregressive sequence length during ASR and TTS inference, lowering computational cost without degrading reconstruction quality. The code is available at https://github.com/facebookresearch/lst.
title Latent Speech-Text Transformer
topic Computation and Language
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2510.06195