MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Mengxue, Diao, Yunfeng, Miao, Changtao, Guo, Zhiqing, Li, Jianshu, Li, Zhe, Zhou, Joey Tianyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918454177562624
author Hu, Mengxue
Diao, Yunfeng
Miao, Changtao
Guo, Zhiqing
Li, Jianshu
Li, Zhe
Zhou, Joey Tianyi
author_facet Hu, Mengxue
Diao, Yunfeng
Miao, Changtao
Guo, Zhiqing
Li, Jianshu
Li, Zhe
Zhou, Joey Tianyi
contents The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality alone, while the few incorporating audio are largely confined to facial deepfakes--a limitation that fails to address the expanding landscape of general multimodal AI-generated content and substantially impedes the development of trustworthy detection systems. To bridge this critical gap, we introduce the Multimodal Video-Audio Dataset (MVAD), the first comprehensive dataset specifically designed for detecting AI-generated multimodal video-audio content. Our dataset exhibits three key characteristics: (1) genuine multimodality with samples generated according to three realistic video-audio forgery patterns; (2) high perceptual quality achieved through diverse state-of-the-art generative models; and (3) comprehensive diversity spanning realistic and anime visual styles, four content categories (humans, animals, objects, and scenes), and four video-audio multimodal data types. Our dataset will be available at https://github.com/HuMengXue0104/MVAD.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
Hu, Mengxue
Diao, Yunfeng
Miao, Changtao
Guo, Zhiqing
Li, Jianshu
Li, Zhe
Zhou, Joey Tianyi
Computer Vision and Pattern Recognition
The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality alone, while the few incorporating audio are largely confined to facial deepfakes--a limitation that fails to address the expanding landscape of general multimodal AI-generated content and substantially impedes the development of trustworthy detection systems. To bridge this critical gap, we introduce the Multimodal Video-Audio Dataset (MVAD), the first comprehensive dataset specifically designed for detecting AI-generated multimodal video-audio content. Our dataset exhibits three key characteristics: (1) genuine multimodality with samples generated according to three realistic video-audio forgery patterns; (2) high perceptual quality achieved through diverse state-of-the-art generative models; and (3) comprehensive diversity spanning realistic and anime visual styles, four content categories (humans, animals, objects, and scenes), and four video-audio multimodal data types. Our dataset will be available at https://github.com/HuMengXue0104/MVAD.
title MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.00336