LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xiaohan, Xiang, Hongyu, Ye, Shengze, Li, Song, Tian, Zhengkun, Chen, Guanyu, Ding, Ke, Wan, Guanglu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908598888562688
author Zhao, Xiaohan
Xiang, Hongyu
Ye, Shengze
Li, Song
Tian, Zhengkun
Chen, Guanyu
Ding, Ke
Wan, Guanglu
author_facet Zhao, Xiaohan
Xiang, Hongyu
Ye, Shengze
Li, Song
Tian, Zhengkun
Chen, Guanyu
Ding, Ke
Wan, Guanglu
contents This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
Zhao, Xiaohan
Xiang, Hongyu
Ye, Shengze
Li, Song
Tian, Zhengkun
Chen, Guanyu
Ding, Ke
Wan, Guanglu
Audio and Speech Processing
Sound
This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.
title LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.15227