STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Tao, Zhao, Zhiyuan, Xie, Yifan, Ye, Yuqi, Luo, Xiangyang, Guan, Xun, Li, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908276863533056
author Feng, Tao
Zhao, Zhiyuan
Xie, Yifan
Ye, Yuqi
Luo, Xiangyang
Guan, Xun
Li, Yu
author_facet Feng, Tao
Zhao, Zhiyuan
Xie, Yifan
Ye, Yuqi
Luo, Xiangyang
Guan, Xun
Li, Yu
contents We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory consumption, this method leverages STFT for compact spectral representation and introduces unwrapped phase derivatives as auxiliary features. Our architecture employs parallel magnitude and phase processing branches enhanced by advanced feature extraction mechanisms. By relaxing strict phase reconstruction constraints while maintaining phase-aware processing, we achieve superior perceptual quality. Experimental results demonstrate that STFTCodec outperforms both waveform-based and spectral-based approaches across multiple bitrates, while offering unique flexibility in compression ratio adjustment through STFT parameter modification without architectural changes.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
Feng, Tao
Zhao, Zhiyuan
Xie, Yifan
Ye, Yuqi
Luo, Xiangyang
Guan, Xun
Li, Yu
Sound
Audio and Speech Processing
We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory consumption, this method leverages STFT for compact spectral representation and introduces unwrapped phase derivatives as auxiliary features. Our architecture employs parallel magnitude and phase processing branches enhanced by advanced feature extraction mechanisms. By relaxing strict phase reconstruction constraints while maintaining phase-aware processing, we achieve superior perceptual quality. Experimental results demonstrate that STFTCodec outperforms both waveform-based and spectral-based approaches across multiple bitrates, while offering unique flexibility in compression ratio adjustment through STFT parameter modification without architectural changes.
title STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.16989