Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Di, Huang, Yihao, Guo, Qing, Juefei-Xu, Felix, Jia, Xiaojun, Wang, Run, Pu, Geguang, Liu, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913366906241024
author Yang, Di
Huang, Yihao
Guo, Qing
Juefei-Xu, Felix
Jia, Xiaojun
Wang, Run
Pu, Geguang
Liu, Yang
author_facet Yang, Di
Huang, Yihao
Guo, Qing
Juefei-Xu, Felix
Jia, Xiaojun
Wang, Run
Pu, Geguang
Liu, Yang
contents The widespread use of diffusion methods enables the creation of highly realistic images on demand, thereby posing significant risks to the integrity and safety of online information and highlighting the necessity of DeepFake detection. Our analysis of features extracted by traditional image encoders reveals that both low-level and high-level features offer distinct advantages in identifying DeepFake images produced by various diffusion methods. Inspired by this finding, we aim to develop an effective representation that captures both low-level and high-level features to detect diffusion-based DeepFakes. To address the problem, we propose a text modality-oriented feature extraction method, termed TOFE. Specifically, for a given target image, the representation we discovered is a corresponding text embedding that can guide the generation of the target image with a specific text-to-image model. Experiments conducted across ten diffusion types demonstrate the efficacy of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake
Yang, Di
Huang, Yihao
Guo, Qing
Juefei-Xu, Felix
Jia, Xiaojun
Wang, Run
Pu, Geguang
Liu, Yang
Computer Vision and Pattern Recognition
The widespread use of diffusion methods enables the creation of highly realistic images on demand, thereby posing significant risks to the integrity and safety of online information and highlighting the necessity of DeepFake detection. Our analysis of features extracted by traditional image encoders reveals that both low-level and high-level features offer distinct advantages in identifying DeepFake images produced by various diffusion methods. Inspired by this finding, we aim to develop an effective representation that captures both low-level and high-level features to detect diffusion-based DeepFakes. To address the problem, we propose a text modality-oriented feature extraction method, termed TOFE. Specifically, for a given target image, the representation we discovered is a corresponding text embedding that can guide the generation of the target image with a specific text-to-image model. Experiments conducted across ten diffusion types demonstrate the efficacy of our proposed method.
title Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18071