Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xuanjun, Cheng, Shih-Peng, Du, Jiawei, Zhang, Lin, Miao, Xiaoxiao, Wang, Chung-Che, Wu, Haibin, Lee, Hung-yi, Jang, Jyh-Shing Roger
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908477703585792
author Chen, Xuanjun
Cheng, Shih-Peng
Du, Jiawei
Zhang, Lin
Miao, Xiaoxiao
Wang, Chung-Che
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
author_facet Chen, Xuanjun
Cheng, Shih-Peng
Du, Jiawei
Zhang, Lin
Miao, Xiaoxiao
Wang, Chung-Che
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
contents Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02000
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
Chen, Xuanjun
Cheng, Shih-Peng
Du, Jiawei
Zhang, Lin
Miao, Xiaoxiao
Wang, Chung-Che
Wu, Haibin
Lee, Hung-yi
Jang, Jyh-Shing Roger
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Image and Video Processing
Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data.
title Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2508.02000