Efficient Interleaved Speech Modeling through Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nouriborji, Mohammadmahdi, Rohanian, Morteza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917031284047872
author Nouriborji, Mohammadmahdi
Rohanian, Morteza
author_facet Nouriborji, Mohammadmahdi
Rohanian, Morteza
contents Current speech language models exceed the size and latency constraints of many deployment environments. We build compact, expressive speech generation models through layer-aligned distillation, matching hidden states, attention maps, and softened logits to compress large multimodal transformers by 3x with minimal loss in performance. We introduce TinyWave, a family of 2B-parameter models for speech-to-speech and interleaved speech-text generation, trained on 50,000 hours of public audio. TinyWave supports (i) speech-only generation using phonetic or expressive tokens and (ii) mixed speech-text continuations. Evaluation on Libri-Light shows TinyWave within 1.4 normalized perplexity points of its teacher. Accuracy on spoken StoryCloze and SALMon reaches 93-97% of the teacher's performance, outperforming size-matched baselines. These models are optimized for deployment on commodity hardware, enabling applications in real-time conversational agents, assistive technologies, and low-resource environments. We release models, training code, and evaluation scripts to support reproducible research on compact, expressive speech generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Interleaved Speech Modeling through Knowledge Distillation
Nouriborji, Mohammadmahdi
Rohanian, Morteza
Sound
Computation and Language
Audio and Speech Processing
Current speech language models exceed the size and latency constraints of many deployment environments. We build compact, expressive speech generation models through layer-aligned distillation, matching hidden states, attention maps, and softened logits to compress large multimodal transformers by 3x with minimal loss in performance. We introduce TinyWave, a family of 2B-parameter models for speech-to-speech and interleaved speech-text generation, trained on 50,000 hours of public audio. TinyWave supports (i) speech-only generation using phonetic or expressive tokens and (ii) mixed speech-text continuations. Evaluation on Libri-Light shows TinyWave within 1.4 normalized perplexity points of its teacher. Accuracy on spoken StoryCloze and SALMon reaches 93-97% of the teacher's performance, outperforming size-matched baselines. These models are optimized for deployment on commodity hardware, enabling applications in real-time conversational agents, assistive technologies, and low-resource environments. We release models, training code, and evaluation scripts to support reproducible research on compact, expressive speech generation.
title Efficient Interleaved Speech Modeling through Knowledge Distillation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.23670