Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahreza, Hatef Otroshi, George, Anjith, Marcel, Sébastien
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915746368454656
author Shahreza, Hatef Otroshi
George, Anjith
Marcel, Sébastien
author_facet Shahreza, Hatef Otroshi
George, Anjith
Marcel, Sébastien
contents Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we conduct a systematic evaluation of state-of-the-art MLLMs for heterogeneous face recognition (HFR), where enrollment and probe images are from different sensing modalities, including visual (VIS), near infrared (NIR), short-wave infrared (SWIR), and thermal camera. We benchmark multiple open-source MLLMs across several cross-modality scenarios, including VIS-NIR, VIS-SWIR, and VIS-THERMAL face recognition. The recognition performance of MLLMs is evaluated using biometric protocols and based on different metrics, including Acquire Rate, Equal Error Rate (EER), and True Accept Rate (TAR). Our results reveal substantial performance gaps between MLLMs and classical face recognition systems, particularly under challenging cross-spectral conditions, in spite of recent advances in MLLMs. Our findings highlight the limitations of current MLLMs for HFR and also the importance of rigorous biometric evaluation when considering their deployment in face recognition systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15406
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition
Shahreza, Hatef Otroshi
George, Anjith
Marcel, Sébastien
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance on a wide range of vision-language tasks, raising interest in their potential use for biometric applications. In this paper, we conduct a systematic evaluation of state-of-the-art MLLMs for heterogeneous face recognition (HFR), where enrollment and probe images are from different sensing modalities, including visual (VIS), near infrared (NIR), short-wave infrared (SWIR), and thermal camera. We benchmark multiple open-source MLLMs across several cross-modality scenarios, including VIS-NIR, VIS-SWIR, and VIS-THERMAL face recognition. The recognition performance of MLLMs is evaluated using biometric protocols and based on different metrics, including Acquire Rate, Equal Error Rate (EER), and True Accept Rate (TAR). Our results reveal substantial performance gaps between MLLMs and classical face recognition systems, particularly under challenging cross-spectral conditions, in spite of recent advances in MLLMs. Our findings highlight the limitations of current MLLMs for HFR and also the importance of rigorous biometric evaluation when considering their deployment in face recognition systems.
title Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.15406