InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Kun-Hsiang, Tseng, Yu-Wen, Huang, Kang-Yang, Wu, Jhih-Ciang, Cheng, Wen-Huang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915417460572160
author Lin, Kun-Hsiang
Tseng, Yu-Wen
Huang, Kang-Yang
Wu, Jhih-Ciang
Cheng, Wen-Huang
author_facet Lin, Kun-Hsiang
Tseng, Yu-Wen
Huang, Kang-Yang
Wu, Jhih-Ciang
Cheng, Wen-Huang
contents Face anti-spoofing (FAS) aims to construct a robust system that can withstand diverse attacks. While recent efforts have concentrated mainly on cross-domain generalization, two significant challenges persist: limited semantic understanding of attack types and training redundancy across domains. We address the first by integrating vision-language models (VLMs) to enhance the perception of visual input. For the second challenge, we employ a meta-domain strategy to learn a unified model that generalizes well across multiple domains. Our proposed InstructFLIP is a novel instruction-tuned framework that leverages VLMs to enhance generalization via textual guidance trained solely on a single domain. At its core, InstructFLIP explicitly decouples instructions into content and style components, where content-based instructions focus on the essential semantics of spoofing, and style-based instructions consider variations related to the environment and camera characteristics. Extensive experiments demonstrate the effectiveness of InstructFLIP by outperforming SOTA models in accuracy and substantially reducing training redundancy across diverse domains in FAS. Project website is available at https://kunkunlin1221.github.io/InstructFLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
Lin, Kun-Hsiang
Tseng, Yu-Wen
Huang, Kang-Yang
Wu, Jhih-Ciang
Cheng, Wen-Huang
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Face anti-spoofing (FAS) aims to construct a robust system that can withstand diverse attacks. While recent efforts have concentrated mainly on cross-domain generalization, two significant challenges persist: limited semantic understanding of attack types and training redundancy across domains. We address the first by integrating vision-language models (VLMs) to enhance the perception of visual input. For the second challenge, we employ a meta-domain strategy to learn a unified model that generalizes well across multiple domains. Our proposed InstructFLIP is a novel instruction-tuned framework that leverages VLMs to enhance generalization via textual guidance trained solely on a single domain. At its core, InstructFLIP explicitly decouples instructions into content and style components, where content-based instructions focus on the essential semantics of spoofing, and style-based instructions consider variations related to the environment and camera characteristics. Extensive experiments demonstrate the effectiveness of InstructFLIP by outperforming SOTA models in accuracy and substantially reducing training redundancy across diverse domains in FAS. Project website is available at https://kunkunlin1221.github.io/InstructFLIP.
title InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2507.12060