FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gavas, Ekta, Banerjee, Sudipta, Hegde, Chinmay, Memon, Nasir
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911587367911424
author Gavas, Ekta
Banerjee, Sudipta
Hegde, Chinmay
Memon, Nasir
author_facet Gavas, Ekta
Banerjee, Sudipta
Hegde, Chinmay
Memon, Nasir
contents Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their ability to analyze biometric data is relatively underexplored. In this work, we investigate the effectiveness of MLLMs in understanding fine structural and textural details present in fingerprint images. To this end, we design a comprehensive benchmark, FPBench, to evaluate 20 MLLMs (open-source and proprietary models) across 7 real and synthetic datasets on a suite of 8 biometric and forensic tasks (e.g., pattern analysis, fingerprint verification, real versus synthetic classification, etc.) using zero-shot and chain-of-thought prompting strategies. We further fine-tune vision and language encoders on a subset of open-source MLLMs to demonstrate domain adaptation. FPBench is a novel benchmark designed as a first step towards developing foundation models in fingerprints. Our findings indicate fine-tuning of vision and language encoders improves the performance by 7%-39%. Our codes are available at https://github.com/Ektagavas/FPBench.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18073
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
Gavas, Ekta
Banerjee, Sudipta
Hegde, Chinmay
Memon, Nasir
Computer Vision and Pattern Recognition
Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their ability to analyze biometric data is relatively underexplored. In this work, we investigate the effectiveness of MLLMs in understanding fine structural and textural details present in fingerprint images. To this end, we design a comprehensive benchmark, FPBench, to evaluate 20 MLLMs (open-source and proprietary models) across 7 real and synthetic datasets on a suite of 8 biometric and forensic tasks (e.g., pattern analysis, fingerprint verification, real versus synthetic classification, etc.) using zero-shot and chain-of-thought prompting strategies. We further fine-tune vision and language encoders on a subset of open-source MLLMs to demonstrate domain adaptation. FPBench is a novel benchmark designed as a first step towards developing foundation models in fingerprints. Our findings indicate fine-tuning of vision and language encoders improves the performance by 7%-39%. Our codes are available at https://github.com/Ektagavas/FPBench.
title FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18073