HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhili Nicholas, Han, Soyeon Caren, Wang, Qizhou, Leckie, Christopher
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908802381512704
author Liang, Zhili Nicholas
Han, Soyeon Caren
Wang, Qizhou
Leckie, Christopher
author_facet Liang, Zhili Nicholas
Han, Soyeon Caren
Wang, Qizhou
Leckie, Christopher
contents Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01032
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
Liang, Zhili Nicholas
Han, Soyeon Caren
Wang, Qizhou
Leckie, Christopher
Sound
Artificial Intelligence
Audio and Speech Processing
Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
title HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2602.01032