Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908477703585792 |
|---|---|
| author | Chen, Xuanjun Cheng, Shih-Peng Du, Jiawei Zhang, Lin Miao, Xiaoxiao Wang, Chung-Che Wu, Haibin Lee, Hung-yi Jang, Jyh-Shing Roger |
| author_facet | Chen, Xuanjun Cheng, Shih-Peng Du, Jiawei Zhang, Lin Miao, Xiaoxiao Wang, Chung-Che Wu, Haibin Lee, Hung-yi Jang, Jyh-Shing Roger |
| contents | Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_02000 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling Chen, Xuanjun Cheng, Shih-Peng Du, Jiawei Zhang, Lin Miao, Xiaoxiao Wang, Chung-Che Wu, Haibin Lee, Hung-yi Jang, Jyh-Shing Roger Sound Computer Vision and Pattern Recognition Audio and Speech Processing Image and Video Processing Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest remaining identical to the original. To tackle this, we propose a Hierarchical Boundary Modeling Network (HBMNet), which includes three modules: an Audio-Visual Feature Encoder that extracts discriminative frame-level representations, a Coarse Proposal Generator that predicts candidate boundary regions, and a Fine-grained Probabilities Generator that refines these proposals using bidirectional boundary-content probabilities. From the modality perspective, we enhance audio-visual learning through dedicated encoding and fusion, reinforced by frame-level supervision to boost discriminability. From the temporal perspective, HBMNet integrates multi-scale cues and bidirectional boundary-content relationships. Experiments show that encoding and fusion primarily improve precision, while frame-level supervision boosts recall. Each module (audio-visual fusion, temporal scales, bi-directionality) contributes complementary benefits, collectively enhancing localization performance. HBMNet outperforms BA-TFD and UMMAFormer and shows improved potential scalability with more training data. |
| title | Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling |
| topic | Sound Computer Vision and Pattern Recognition Audio and Speech Processing Image and Video Processing |
| url | https://arxiv.org/abs/2508.02000 |