Salvato in:
Dettagli Bibliografici
Autori principali: Sugiura, Issa, Kurita, Shuhei, Oda, Yusuke, Higashinaka, Ryuichiro
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2509.14882
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915836235612160
author Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Higashinaka, Ryuichiro
author_facet Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Higashinaka, Ryuichiro
contents Speech Language Models (SpeechLMs) model tokenized speech to capture both semantic and acoustic information. When neural audio codecs based on Residual Vector Quantization (RVQ) are used as audio tokenizers, they produce multiple discrete tokens per time step, yielding inherently multi-level representations. To process these multi-level tokens together, prior work typically adopts hierarchical architectures to capture this structure. In contrast, recent progress in NLP has progressively reduced architectural inductive biases, moving toward simpler and more scalable single-Transformer architectures. In this work, we propose Llama-Mimi, which flattens multi-level RVQ tokens produced by the Mimi neural audio codec into a single sequence and models them autoregressively with a Transformer decoder. We show that Llama-Mimi outperforms a CSM-based hierarchical model on most tasks and achieves the best performance on acoustic consistency. Our models, code, and speech samples are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14882
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
Sugiura, Issa
Kurita, Shuhei
Oda, Yusuke
Higashinaka, Ryuichiro
Computation and Language
Speech Language Models (SpeechLMs) model tokenized speech to capture both semantic and acoustic information. When neural audio codecs based on Residual Vector Quantization (RVQ) are used as audio tokenizers, they produce multiple discrete tokens per time step, yielding inherently multi-level representations. To process these multi-level tokens together, prior work typically adopts hierarchical architectures to capture this structure. In contrast, recent progress in NLP has progressively reduced architectural inductive biases, moving toward simpler and more scalable single-Transformer architectures. In this work, we propose Llama-Mimi, which flattens multi-level RVQ tokens produced by the Mimi neural audio codec into a single sequence and models them autoregressively with a Transformer decoder. We show that Llama-Mimi outperforms a CSM-based hierarchical model on most tasks and achieves the best performance on acoustic consistency. Our models, code, and speech samples are publicly available.
title Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
topic Computation and Language
url https://arxiv.org/abs/2509.14882