Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914404370481152 |
|---|---|
| author | Wu, Jinyang Pan, Zihan Zhang, Qiquan Bhupendra, Sailor Hardik Mondal, Soumik |
| author_facet | Wu, Jinyang Pan, Zihan Zhang, Qiquan Bhupendra, Sailor Hardik Mondal, Soumik |
| contents | Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_16914 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection Wu, Jinyang Pan, Zihan Zhang, Qiquan Bhupendra, Sailor Hardik Mondal, Soumik Sound Artificial Intelligence Computation and Language Audio and Speech Processing Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines. |
| title | Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection |
| topic | Sound Artificial Intelligence Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2603.16914 |