Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hanglei, Guo, Yiwei, Li, Zhihan, Hao, Xiang, Chen, Xie, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915483702263808
author Zhang, Hanglei
Guo, Yiwei
Li, Zhihan
Hao, Xiang
Chen, Xie
Yu, Kai
author_facet Zhang, Hanglei
Guo, Yiwei
Li, Zhihan
Hao, Xiang
Chen, Xie
Yu, Kai
contents Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
Zhang, Hanglei
Guo, Yiwei
Li, Zhihan
Hao, Xiang
Chen, Xie
Yu, Kai
Audio and Speech Processing
Artificial Intelligence
Sound
Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.
title Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2505.16845