Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Kaiqing, Lin, Yuzhen, Li, Weixiang, Yao, Taiping, Li, Bin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910908638298112
author Lin, Kaiqing
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Li, Bin
author_facet Lin, Kaiqing
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Li, Bin
contents The proliferation of deepfake faces poses huge potential negative impacts on our daily lives. Despite substantial advancements in deepfake detection over these years, the generalizability of existing methods against forgeries from unseen datasets or created by emerging generative models remains constrained. In this paper, inspired by the zero-shot advantages of Vision-Language Models (VLMs), we propose a novel approach that repurposes a well-trained VLM for general deepfake detection. Motivated by the model reprogramming paradigm that manipulates the model prediction via input perturbations, our method can reprogram a pre-trained VLM model (e.g., CLIP) solely based on manipulating its input without tuning the inner parameters. First, learnable visual perturbations are used to refine feature extraction for deepfake detection. Then, we exploit information of face embedding to create sample-level adaptative text prompts, improving the performance. Extensive experiments on several popular benchmark datasets demonstrate that (1) the cross-dataset and cross-manipulation performances of deepfake detection can be significantly and consistently improved (e.g., over 88\% AUC in cross-dataset setting from FF++ to WildDeepfake); (2) the superior performances are achieved with fewer trainable parameters, making it a promising approach for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02664
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
Lin, Kaiqing
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Li, Bin
Computer Vision and Pattern Recognition
The proliferation of deepfake faces poses huge potential negative impacts on our daily lives. Despite substantial advancements in deepfake detection over these years, the generalizability of existing methods against forgeries from unseen datasets or created by emerging generative models remains constrained. In this paper, inspired by the zero-shot advantages of Vision-Language Models (VLMs), we propose a novel approach that repurposes a well-trained VLM for general deepfake detection. Motivated by the model reprogramming paradigm that manipulates the model prediction via input perturbations, our method can reprogram a pre-trained VLM model (e.g., CLIP) solely based on manipulating its input without tuning the inner parameters. First, learnable visual perturbations are used to refine feature extraction for deepfake detection. Then, we exploit information of face embedding to create sample-level adaptative text prompts, improving the performance. Extensive experiments on several popular benchmark datasets demonstrate that (1) the cross-dataset and cross-manipulation performances of deepfake detection can be significantly and consistently improved (e.g., over 88\% AUC in cross-dataset setting from FF++ to WildDeepfake); (2) the superior performances are achieved with fewer trainable parameters, making it a promising approach for real-world applications.
title Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.02664