Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Peipeng, Fei, Jianwei, Gao, Hui, Feng, Xuan, Xia, Zhihua, Chang, Chip Hong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918049458683904
author Yu, Peipeng
Fei, Jianwei
Gao, Hui
Feng, Xuan
Xia, Zhihua
Chang, Chip Hong
author_facet Yu, Peipeng
Fei, Jianwei
Gao, Hui
Feng, Xuan
Xia, Zhihua
Chang, Chip Hong
contents Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs' potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
Yu, Peipeng
Fei, Jianwei
Gao, Hui
Feng, Xuan
Xia, Zhihua
Chang, Chip Hong
Computer Vision and Pattern Recognition
Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs' potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities.
title Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.14853