Saved in:
Bibliographic Details
Main Authors: Wu, Xuecheng, Huang, Danlei, Sun, Heli, Yin, Xinyi, Wang, Yifan, Wang, Hao, Zhang, Jia, Wang, Fei, Guo, Peihao, Xing, Suyu, Xue, Junxiao, He, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.22781
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918107816132608
author Wu, Xuecheng
Huang, Danlei
Sun, Heli
Yin, Xinyi
Wang, Yifan
Wang, Hao
Zhang, Jia
Wang, Fei
Guo, Peihao
Xing, Suyu
Xue, Junxiao
He, Liang
author_facet Wu, Xuecheng
Huang, Danlei
Sun, Heli
Yin, Xinyi
Wang, Yifan
Wang, Hao
Zhang, Jia
Wang, Fei
Guo, Peihao
Xing, Suyu
Xue, Junxiao
He, Liang
contents Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
Wu, Xuecheng
Huang, Danlei
Sun, Heli
Yin, Xinyi
Wang, Yifan
Wang, Hao
Zhang, Jia
Wang, Fei
Guo, Peihao
Xing, Suyu
Xue, Junxiao
He, Liang
Computer Vision and Pattern Recognition
Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set.
title HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.22781