Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jinyang, Pan, Zihan, Zhang, Qiquan, Bhupendra, Sailor Hardik, Mondal, Soumik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914404370481152
author Wu, Jinyang
Pan, Zihan
Zhang, Qiquan
Bhupendra, Sailor Hardik
Mondal, Soumik
author_facet Wu, Jinyang
Pan, Zihan
Zhang, Qiquan
Bhupendra, Sailor Hardik
Mondal, Soumik
contents Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
Wu, Jinyang
Pan, Zihan
Zhang, Qiquan
Bhupendra, Sailor Hardik
Mondal, Soumik
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines.
title Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2603.16914