CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Ajian, Xue, Shuai, Gan, Jianwen, Wan, Jun, Liang, Yanyan, Deng, Jiankang, Escalera, Sergio, Lei, Zhen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929284578279424
author Liu, Ajian
Xue, Shuai
Gan, Jianwen
Wan, Jun
Liang, Yanyan
Deng, Jiankang
Escalera, Sergio
Lei, Zhen
author_facet Liu, Ajian
Xue, Shuai
Gan, Jianwen
Wan, Jun
Liang, Yanyan
Deng, Jiankang
Escalera, Sergio
Lei, Zhen
contents Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14333
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing
Liu, Ajian
Xue, Shuai
Gan, Jianwen
Wan, Jun
Liang, Yanyan
Deng, Jiankang
Escalera, Sergio
Lei, Zhen
Computer Vision and Pattern Recognition
Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets.
title CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.14333