Benchmarking Multimodal Large Language Models for Face Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahreza, Hatef Otroshi, Marcel, Sébastien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908598002515968
author Shahreza, Hatef Otroshi
Marcel, Sébastien
author_facet Shahreza, Hatef Otroshi
Marcel, Sébastien
contents Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance of open-source MLLMs needs to be evaluated and compared with existing face recognition models on standard benchmarks with similar protocol. In this work, we present a systematic benchmark of state-of-the-art MLLMs for face recognition on several face recognition datasets, including LFW, CALFW, CPLFW, CFP, AgeDB and RFW. Experimental results reveal that while MLLMs capture rich semantic cues useful for face-related tasks, they lag behind specialized models in high-precision recognition scenarios in zero-shot applications. This benchmark provides a foundation for advancing MLLM-based face recognition, offering insights for the design of next-generation models with higher accuracy and generalization. The source code of our benchmark is publicly available in the project page.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Multimodal Large Language Models for Face Recognition
Shahreza, Hatef Otroshi
Marcel, Sébastien
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance of open-source MLLMs needs to be evaluated and compared with existing face recognition models on standard benchmarks with similar protocol. In this work, we present a systematic benchmark of state-of-the-art MLLMs for face recognition on several face recognition datasets, including LFW, CALFW, CPLFW, CFP, AgeDB and RFW. Experimental results reveal that while MLLMs capture rich semantic cues useful for face-related tasks, they lag behind specialized models in high-precision recognition scenarios in zero-shot applications. This benchmark provides a foundation for advancing MLLM-based face recognition, offering insights for the design of next-generation models with higher accuracy and generalization. The source code of our benchmark is publicly available in the project page.
title Benchmarking Multimodal Large Language Models for Face Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.14866