MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Croitoru, Florinel-Alin, Hondru, Vlad, Popescu, Marius, Ionescu, Radu Tudor, Khan, Fahad Shahbaz, Shah, Mubarak
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916739858563072
author Croitoru, Florinel-Alin
Hondru, Vlad
Popescu, Marius
Ionescu, Radu Tudor
Khan, Fahad Shahbaz
Shah, Mubarak
author_facet Croitoru, Florinel-Alin
Hondru, Vlad
Popescu, Marius
Ionescu, Radu Tudor
Khan, Fahad Shahbaz
Shah, Mubarak
contents We present the first large-scale open-set benchmark for multilingual audio-video deepfake detection. Our dataset comprises over 250 hours of real and fake videos across eight languages, with 60% of data being generated. For each language, the fake videos are generated with seven distinct deepfake generation models, selected based on the quality of the generated content. We organize the training, validation and test splits such that only a subset of the chosen generative models and languages are available during training, thus creating several challenging open-set evaluation setups. We perform experiments with various pre-trained and fine-tuned deepfake detectors proposed in recent literature. Our results show that state-of-the-art detectors are not currently able to maintain their performance levels when tested in our open-set scenarios. We publicly release our data and code at: https://huggingface.co/datasets/unibuc-cs/MAVOS-DD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
Croitoru, Florinel-Alin
Hondru, Vlad
Popescu, Marius
Ionescu, Radu Tudor
Khan, Fahad Shahbaz
Shah, Mubarak
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
We present the first large-scale open-set benchmark for multilingual audio-video deepfake detection. Our dataset comprises over 250 hours of real and fake videos across eight languages, with 60% of data being generated. For each language, the fake videos are generated with seven distinct deepfake generation models, selected based on the quality of the generated content. We organize the training, validation and test splits such that only a subset of the chosen generative models and languages are available during training, thus creating several challenging open-set evaluation setups. We perform experiments with various pre-trained and fine-tuned deepfake detectors proposed in recent literature. Our results show that state-of-the-art detectors are not currently able to maintain their performance levels when tested in our open-set scenarios. We publicly release our data and code at: https://huggingface.co/datasets/unibuc-cs/MAVOS-DD.
title MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2505.11109