PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Yang, Fu, Haitao, Chen, Chuankai, Li, Zida, Zhang, Haoyu, Zhao, Jianjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909203103219712
author Hou, Yang
Fu, Haitao
Chen, Chuankai
Li, Zida
Zhang, Haoyu
Zhao, Jianjun
author_facet Hou, Yang
Fu, Haitao
Chen, Chuankai
Li, Zida
Zhang, Haoyu
Zhao, Jianjun
contents With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these growing threats. However, as a key factor in training and validating deepfake detectors, most existing deepfake datasets primarily focus on the visual modal, and the few that are multimodal employ outdated techniques, and their audio content is limited to a single language, thereby failing to represent the cutting-edge advancements and globalization trends in current deepfake technologies. To address this gap, we propose a novel, multilingual, and multimodal deepfake dataset: PolyGlotFake. It includes content in seven languages, created using a variety of cutting-edge and popular Text-to-Speech, voice cloning, and lip-sync technologies. We conduct comprehensive experiments using state-of-the-art detection methods on PolyGlotFake dataset. These experiments demonstrate the dataset's significant challenges and its practical value in advancing research into multimodal deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
Hou, Yang
Fu, Haitao
Chen, Chuankai
Li, Zida
Zhang, Haoyu
Zhao, Jianjun
Sound
Artificial Intelligence
Audio and Speech Processing
68T45
I.4.9
With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these growing threats. However, as a key factor in training and validating deepfake detectors, most existing deepfake datasets primarily focus on the visual modal, and the few that are multimodal employ outdated techniques, and their audio content is limited to a single language, thereby failing to represent the cutting-edge advancements and globalization trends in current deepfake technologies. To address this gap, we propose a novel, multilingual, and multimodal deepfake dataset: PolyGlotFake. It includes content in seven languages, created using a variety of cutting-edge and popular Text-to-Speech, voice cloning, and lip-sync technologies. We conduct comprehensive experiments using state-of-the-art detection methods on PolyGlotFake dataset. These experiments demonstrate the dataset's significant challenges and its practical value in advancing research into multimodal deepfake detection.
title PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
topic Sound
Artificial Intelligence
Audio and Speech Processing
68T45
I.4.9
url https://arxiv.org/abs/2405.08838