TS3-Codec: Transformer-Based Simple Streaming Single Codec

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Haibin, Kanda, Naoyuki, Eskimez, Sefik Emre, Li, Jinyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912348471558144
author Wu, Haibin
Kanda, Naoyuki
Eskimez, Sefik Emre
Li, Jinyu
author_facet Wu, Haibin
Kanda, Naoyuki
Eskimez, Sefik Emre
Li, Jinyu
contents Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convolution-based, the performance of NACs with a purely transformer-based, and convolution-free architecture remains unexplored. This paper introduces TS3-Codec, a Transformer-Based Simple Streaming Single Codec. TS3-Codec consists of only a stack of transformer layers with a few linear layers, offering greater simplicity and expressiveness by fully eliminating convolution layers that require careful hyperparameter tuning and large computations. Under the streaming setup, the proposed TS3-Codec achieves comparable or superior performance compared to the codec with state-of-the-art convolution-based architecture while requiring only 12% of the computation and 77% of bitrate. Furthermore, it significantly outperforms the convolution-based codec when using similar computational resources.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TS3-Codec: Transformer-Based Simple Streaming Single Codec
Wu, Haibin
Kanda, Naoyuki
Eskimez, Sefik Emre
Li, Jinyu
Audio and Speech Processing
Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convolution-based, the performance of NACs with a purely transformer-based, and convolution-free architecture remains unexplored. This paper introduces TS3-Codec, a Transformer-Based Simple Streaming Single Codec. TS3-Codec consists of only a stack of transformer layers with a few linear layers, offering greater simplicity and expressiveness by fully eliminating convolution layers that require careful hyperparameter tuning and large computations. Under the streaming setup, the proposed TS3-Codec achieves comparable or superior performance compared to the codec with state-of-the-art convolution-based architecture while requiring only 12% of the computation and 77% of bitrate. Furthermore, it significantly outperforms the convolution-based codec when using similar computational resources.
title TS3-Codec: Transformer-Based Simple Streaming Single Codec
topic Audio and Speech Processing
url https://arxiv.org/abs/2411.18803