How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahzad, Sahibzada Adil, Hashmi, Ammarah, Peng, Yan-Tsung, Tsao, Yu, Wang, Hsin-Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909389273694208
author Shahzad, Sahibzada Adil
Hashmi, Ammarah
Peng, Yan-Tsung
Tsao, Yu
Wang, Hsin-Min
author_facet Shahzad, Sahibzada Adil
Hashmi, Ammarah
Peng, Yan-Tsung
Tsao, Yu
Wang, Hsin-Min
contents Multimodal deepfakes involving audiovisual manipulations are a growing threat because they are difficult to detect with the naked eye or using unimodal deep learningbased forgery detection methods. Audiovisual forensic models, while more capable than unimodal models, require large training datasets and are computationally expensive for training and inference. Furthermore, these models lack interpretability and often do not generalize well to unseen manipulations. In this study, we examine the detection capabilities of a large language model (LLM) (i.e., ChatGPT) to identify and account for any possible visual and auditory artifacts and manipulations in audiovisual deepfake content. Extensive experiments are conducted on videos from a benchmark multimodal deepfake dataset to evaluate the detection performance of ChatGPT and compare it with the detection capabilities of state-of-the-art multimodal forensic models and humans. Experimental results demonstrate the importance of domain knowledge and prompt engineering for video forgery detection tasks using LLMs. Unlike approaches based on end-to-end learning, ChatGPT can account for spatial and spatiotemporal artifacts and inconsistencies that may exist within or across modalities. Additionally, we discuss the limitations of ChatGPT for multimedia forensic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09266
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
Shahzad, Sahibzada Adil
Hashmi, Ammarah
Peng, Yan-Tsung
Tsao, Yu
Wang, Hsin-Min
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multimedia
Multimodal deepfakes involving audiovisual manipulations are a growing threat because they are difficult to detect with the naked eye or using unimodal deep learningbased forgery detection methods. Audiovisual forensic models, while more capable than unimodal models, require large training datasets and are computationally expensive for training and inference. Furthermore, these models lack interpretability and often do not generalize well to unseen manipulations. In this study, we examine the detection capabilities of a large language model (LLM) (i.e., ChatGPT) to identify and account for any possible visual and auditory artifacts and manipulations in audiovisual deepfake content. Extensive experiments are conducted on videos from a benchmark multimodal deepfake dataset to evaluate the detection performance of ChatGPT and compare it with the detection capabilities of state-of-the-art multimodal forensic models and humans. Experimental results demonstrate the importance of domain knowledge and prompt engineering for video forgery detection tasks using LLMs. Unlike approaches based on end-to-end learning, ChatGPT can account for spatial and spatiotemporal artifacts and inconsistencies that may exist within or across modalities. Additionally, we discuss the limitations of ChatGPT for multimedia forensic tasks.
title How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multimedia
url https://arxiv.org/abs/2411.09266