Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913366906241024 |
|---|---|
| author | Yang, Di Huang, Yihao Guo, Qing Juefei-Xu, Felix Jia, Xiaojun Wang, Run Pu, Geguang Liu, Yang |
| author_facet | Yang, Di Huang, Yihao Guo, Qing Juefei-Xu, Felix Jia, Xiaojun Wang, Run Pu, Geguang Liu, Yang |
| contents | The widespread use of diffusion methods enables the creation of highly realistic images on demand, thereby posing significant risks to the integrity and safety of online information and highlighting the necessity of DeepFake detection. Our analysis of features extracted by traditional image encoders reveals that both low-level and high-level features offer distinct advantages in identifying DeepFake images produced by various diffusion methods. Inspired by this finding, we aim to develop an effective representation that captures both low-level and high-level features to detect diffusion-based DeepFakes. To address the problem, we propose a text modality-oriented feature extraction method, termed TOFE. Specifically, for a given target image, the representation we discovered is a corresponding text embedding that can guide the generation of the target image with a specific text-to-image model. Experiments conducted across ten diffusion types demonstrate the efficacy of our proposed method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_18071 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake Yang, Di Huang, Yihao Guo, Qing Juefei-Xu, Felix Jia, Xiaojun Wang, Run Pu, Geguang Liu, Yang Computer Vision and Pattern Recognition The widespread use of diffusion methods enables the creation of highly realistic images on demand, thereby posing significant risks to the integrity and safety of online information and highlighting the necessity of DeepFake detection. Our analysis of features extracted by traditional image encoders reveals that both low-level and high-level features offer distinct advantages in identifying DeepFake images produced by various diffusion methods. Inspired by this finding, we aim to develop an effective representation that captures both low-level and high-level features to detect diffusion-based DeepFakes. To address the problem, we propose a text modality-oriented feature extraction method, termed TOFE. Specifically, for a given target image, the representation we discovered is a corresponding text embedding that can guide the generation of the target image with a specific text-to-image model. Experiments conducted across ten diffusion types demonstrate the efficacy of our proposed method. |
| title | Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.18071 |