Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Kaiqing, Yan, Zhiyuan, Zhang, Ke-Yue, Hao, Li, Zhou, Yue, Lin, Yuzhen, Li, Weixiang, Yao, Taiping, Ding, Shouhong, Li, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908689222336512
author Lin, Kaiqing
Yan, Zhiyuan
Zhang, Ke-Yue
Hao, Li
Zhou, Yue
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Ding, Shouhong
Li, Bin
author_facet Lin, Kaiqing
Yan, Zhiyuan
Zhang, Ke-Yue
Hao, Li
Zhou, Yue
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Ding, Shouhong
Li, Bin
contents Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore the valuable prior knowledge of known facial identities, e.g., "VIP individuals" whose authentic facial data are already available. In this paper, we propose \textbf{VIPGuard}, a unified multimodal framework designed to capture fine-grained and comprehensive facial representations of a given identity, compare them against potentially fake or similar-looking faces, and reason over these comparisons to make accurate and explainable predictions. Specifically, our framework consists of three main stages. First, fine-tune a multimodal large language model (MLLM) to learn detailed and structural facial attributes. Second, we perform identity-level discriminative learning to enable the model to distinguish subtle differences between highly similar faces, including real and fake variations. Finally, we introduce user-specific customization, where we model the unique characteristics of the target face identity and perform semantic reasoning via MLLM to enable personalized and explainable deepfake detection. Our framework shows clear advantages over previous detection works, where traditional detectors mainly rely on low-level visual cues and provide no human-understandable explanations, while other MLLM-based models often lack a detailed understanding of specific face identities. To facilitate the evaluation of our method, we built a comprehensive identity-aware benchmark called \textbf{VIPBench} for personalized deepfake detection, involving the latest 7 face-swapping and 7 entire face synthesis techniques for generation. The code is available at https://github.com/KQL11/VIPGuard .
format Preprint
id arxiv_https___arxiv_org_abs_2505_19582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
Lin, Kaiqing
Yan, Zhiyuan
Zhang, Ke-Yue
Hao, Li
Zhou, Yue
Lin, Yuzhen
Li, Weixiang
Yao, Taiping
Ding, Shouhong
Li, Bin
Computer Vision and Pattern Recognition
Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore the valuable prior knowledge of known facial identities, e.g., "VIP individuals" whose authentic facial data are already available. In this paper, we propose \textbf{VIPGuard}, a unified multimodal framework designed to capture fine-grained and comprehensive facial representations of a given identity, compare them against potentially fake or similar-looking faces, and reason over these comparisons to make accurate and explainable predictions. Specifically, our framework consists of three main stages. First, fine-tune a multimodal large language model (MLLM) to learn detailed and structural facial attributes. Second, we perform identity-level discriminative learning to enable the model to distinguish subtle differences between highly similar faces, including real and fake variations. Finally, we introduce user-specific customization, where we model the unique characteristics of the target face identity and perform semantic reasoning via MLLM to enable personalized and explainable deepfake detection. Our framework shows clear advantages over previous detection works, where traditional detectors mainly rely on low-level visual cues and provide no human-understandable explanations, while other MLLM-based models often lack a detailed understanding of specific face identities. To facilitate the evaluation of our method, we built a comprehensive identity-aware benchmark called \textbf{VIPBench} for personalized deepfake detection, involving the latest 7 face-swapping and 7 entire face synthesis techniques for generation. The code is available at https://github.com/KQL11/VIPGuard .
title Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.19582