Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Parker, Julian D, Smirnov, Anton, Pons, Jordi, Carr, CJ, Zukowski, Zack, Evans, Zach, Liu, Xubo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913591508074496
author Parker, Julian D
Smirnov, Anton
Pons, Jordi
Carr, CJ
Zukowski, Zack
Evans, Zach
Liu, Xubo
author_facet Parker, Julian D
Smirnov, Anton
Pons, Jordi
Carr, CJ
Zukowski, Zack
Evans, Zach
Liu, Xubo
contents The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by scaling a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of $400$ or $700$ bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Transformers for Low-Bitrate High-Quality Speech Coding
Parker, Julian D
Smirnov, Anton
Pons, Jordi
Carr, CJ
Zukowski, Zack
Evans, Zach
Liu, Xubo
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Signal Processing
The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by scaling a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of $400$ or $700$ bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.
title Scaling Transformers for Low-Bitrate High-Quality Speech Coding
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2411.19842