Towards General Visual-Linguistic Face Forgery Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Ke, Chen, Shen, Yao, Taiping, Yang, Haozhe, Sun, Xiaoshuai, Ding, Shouhong, Ji, Rongrong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911772394389504
author Sun, Ke
Chen, Shen
Yao, Taiping
Yang, Haozhe
Sun, Xiaoshuai
Ding, Shouhong
Ji, Rongrong
author_facet Sun, Ke
Chen, Shen
Yao, Taiping
Yang, Haozhe
Sun, Xiaoshuai
Ding, Shouhong
Ji, Rongrong
contents Deepfakes are realistic face manipulations that can pose serious threats to security, privacy, and trust. Existing methods mostly treat this task as binary classification, which uses digital labels or mask signals to train the detection model. We argue that such supervisions lack semantic information and interpretability. To address this issues, in this paper, we propose a novel paradigm named Visual-Linguistic Face Forgery Detection(VLFFD), which uses fine-grained sentence-level prompts as the annotation. Since text annotations are not available in current deepfakes datasets, VLFFD first generates the mixed forgery image with corresponding fine-grained prompts via Prompt Forgery Image Generator (PFIG). Then, the fine-grained mixed data and coarse-grained original data and is jointly trained with the Coarse-and-Fine Co-training framework (C2F), enabling the model to gain more generalization and interpretability. The experiments show the proposed method improves the existing detection models on several challenging benchmarks. Furthermore, we have integrated our method with multimodal large models, achieving noteworthy results that demonstrate the potential of our approach. This integration not only enhances the performance of our VLFFD paradigm but also underscores the versatility and adaptability of our method when combined with advanced multimodal technologies, highlighting its potential in tackling the evolving challenges of deepfake detection.
format Preprint
id arxiv_https___arxiv_org_abs_2307_16545
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards General Visual-Linguistic Face Forgery Detection
Sun, Ke
Chen, Shen
Yao, Taiping
Yang, Haozhe
Sun, Xiaoshuai
Ding, Shouhong
Ji, Rongrong
Computer Vision and Pattern Recognition
Deepfakes are realistic face manipulations that can pose serious threats to security, privacy, and trust. Existing methods mostly treat this task as binary classification, which uses digital labels or mask signals to train the detection model. We argue that such supervisions lack semantic information and interpretability. To address this issues, in this paper, we propose a novel paradigm named Visual-Linguistic Face Forgery Detection(VLFFD), which uses fine-grained sentence-level prompts as the annotation. Since text annotations are not available in current deepfakes datasets, VLFFD first generates the mixed forgery image with corresponding fine-grained prompts via Prompt Forgery Image Generator (PFIG). Then, the fine-grained mixed data and coarse-grained original data and is jointly trained with the Coarse-and-Fine Co-training framework (C2F), enabling the model to gain more generalization and interpretability. The experiments show the proposed method improves the existing detection models on several challenging benchmarks. Furthermore, we have integrated our method with multimodal large models, achieving noteworthy results that demonstrate the potential of our approach. This integration not only enhances the performance of our VLFFD paradigm but also underscores the versatility and adaptability of our method when combined with advanced multimodal technologies, highlighting its potential in tackling the evolving challenges of deepfake detection.
title Towards General Visual-Linguistic Face Forgery Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2307.16545