Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Patwardhan, Sushrut, Ramachandra, Raghavendra, Venkatesh, Sushma
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912537242501120
author Patwardhan, Sushrut
Ramachandra, Raghavendra
Venkatesh, Sushma
author_facet Patwardhan, Sushrut
Ramachandra, Raghavendra
Venkatesh, Sushma
contents Morphing attack detection has become an essential component of face recognition systems for ensuring a reliable verification scenario. In this paper, we present a multimodal learning approach that can provide a textual description of morphing attack detection. We first show that zero-shot evaluation of the proposed framework using Contrastive Language-Image Pretraining (CLIP) can yield not only generalizable morphing attack detection, but also predict the most relevant text snippet. We present an extensive analysis of ten different textual prompts that include both short and long textual prompts. These prompts are engineered by considering the human understandable textual snippet. Extensive experiments were performed on a face morphing dataset that was developed using a publicly available face biometric dataset. We present an evaluation of SOTA pre-trained neural networks together with the proposed framework in the zero-shot evaluation of five different morphing generation techniques that are captured in three different mediums.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10110
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
Patwardhan, Sushrut
Ramachandra, Raghavendra
Venkatesh, Sushma
Computer Vision and Pattern Recognition
Artificial Intelligence
Morphing attack detection has become an essential component of face recognition systems for ensuring a reliable verification scenario. In this paper, we present a multimodal learning approach that can provide a textual description of morphing attack detection. We first show that zero-shot evaluation of the proposed framework using Contrastive Language-Image Pretraining (CLIP) can yield not only generalizable morphing attack detection, but also predict the most relevant text snippet. We present an extensive analysis of ten different textual prompts that include both short and long textual prompts. These prompts are engineered by considering the human understandable textual snippet. Extensive experiments were performed on a face morphing dataset that was developed using a publicly available face biometric dataset. We present an evaluation of SOTA pre-trained neural networks together with the proposed framework in the zero-shot evaluation of five different morphing generation techniques that are captured in three different mediums.
title Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.10110