SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Yichen, Gao, Yuhao, Lai, Yingxin, Wang, Hongyang, Feng, Jun, He, Lei, Wan, Jun, Chen, Changsheng, Yu, Zitong, Cao, Xiaochun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913799921991680
author Shi, Yichen
Gao, Yuhao
Lai, Yingxin
Wang, Hongyang
Feng, Jun
He, Lei
Wan, Jun
Chen, Changsheng
Yu, Zitong
Cao, Xiaochun
author_facet Shi, Yichen
Gao, Yuhao
Lai, Yingxin
Wang, Hongyang
Feng, Jun
He, Lei
Wan, Jun
Chen, Changsheng
Yu, Zitong
Cao, Xiaochun
contents Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing and forgery clues in face attack detection tasks remains underexplored. In this paper, we introduce a benchmark, SHIELD, to evaluate MLLMs for face spoofing and forgery detection. Specifically, we design true/false and multiple-choice questions to assess MLLM performance on multimodal face data across two tasks. For the face anti-spoofing task, we evaluate three modalities (i.e., RGB, infrared, and depth) under six attack types. For the face forgery detection task, we evaluate GAN-based and diffusion-based data, incorporating visual and acoustic modalities. We conduct zero-shot and few-shot evaluations in standard and chain of thought (COT) settings. Additionally, we propose a novel multi-attribute chain of thought (MA-COT) paradigm for describing and judging various task-specific and task-irrelevant attributes of face images. The findings of this study demonstrate that MLLMs exhibit strong potential for addressing the challenges associated with the security of facial recognition technology applications.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models
Shi, Yichen
Gao, Yuhao
Lai, Yingxin
Wang, Hongyang
Feng, Jun
He, Lei
Wan, Jun
Chen, Changsheng
Yu, Zitong
Cao, Xiaochun
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing and forgery clues in face attack detection tasks remains underexplored. In this paper, we introduce a benchmark, SHIELD, to evaluate MLLMs for face spoofing and forgery detection. Specifically, we design true/false and multiple-choice questions to assess MLLM performance on multimodal face data across two tasks. For the face anti-spoofing task, we evaluate three modalities (i.e., RGB, infrared, and depth) under six attack types. For the face forgery detection task, we evaluate GAN-based and diffusion-based data, incorporating visual and acoustic modalities. We conduct zero-shot and few-shot evaluations in standard and chain of thought (COT) settings. Additionally, we propose a novel multi-attribute chain of thought (MA-COT) paradigm for describing and judging various task-specific and task-irrelevant attributes of face images. The findings of this study demonstrate that MLLMs exhibit strong potential for addressing the challenges associated with the security of facial recognition technology applications.
title SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.04178