AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Zhixi, Kuckreja, Kartik, Ghosh, Shreya, Chuchra, Akanksha, Khan, Muhammad Haris, Tariq, Usman, Gedeon, Tom, Dhall, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912505035489280
author Cai, Zhixi
Kuckreja, Kartik
Ghosh, Shreya
Chuchra, Akanksha
Khan, Muhammad Haris
Tariq, Usman
Gedeon, Tom
Dhall, Abhinav
author_facet Cai, Zhixi
Kuckreja, Kartik
Ghosh, Shreya
Chuchra, Akanksha
Khan, Muhammad Haris
Tariq, Usman
Gedeon, Tom
Dhall, Abhinav
contents The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
Cai, Zhixi
Kuckreja, Kartik
Ghosh, Shreya
Chuchra, Akanksha
Khan, Muhammad Haris
Tariq, Usman
Gedeon, Tom
Dhall, Abhinav
Computer Vision and Pattern Recognition
The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.
title AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20579