MLLM-based Textual Explanations for Face Comparison
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sony, Redwan, Jain, Anil K, Ross, Arun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
A Parametric Approach to Adversarial Augmentation for Cross-Domain Iris Presentation Attack Detection
von: Pal, Debasmita, et al.
Veröffentlicht: (2024)
von: Pal, Debasmita, et al.
Veröffentlicht: (2024)
Adversarial Watermarking for Face Recognition
von: Yao, Yuguang, et al.
Veröffentlicht: (2024)
von: Yao, Yuguang, et al.
Veröffentlicht: (2024)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
von: Eleftheriadis, Thomas, et al.
Veröffentlicht: (2025)
von: Eleftheriadis, Thomas, et al.
Veröffentlicht: (2025)
GenPalm: Contactless Palmprint Generation with Diffusion Models
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
Universal Fingerprint Generation: Controllable Diffusion Model with Multimodal Conditions
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
Towards A Comprehensive Visual Saliency Explanation Framework for AI-based Face Recognition Systems
von: Lu, Yuhang, et al.
Veröffentlicht: (2024)
von: Lu, Yuhang, et al.
Veröffentlicht: (2024)
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2026)
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
On Missing Scores in Evolving Multibiometric Systems
von: Dale, Melissa R, et al.
Veröffentlicht: (2024)
von: Dale, Melissa R, et al.
Veröffentlicht: (2024)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
FreeVA: Offline MLLM as Training-Free Video Assistant
von: Wu, Wenhao
Veröffentlicht: (2024)
von: Wu, Wenhao
Veröffentlicht: (2024)
Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring
von: Aghakishiyeva, Günel, et al.
Veröffentlicht: (2025)
von: Aghakishiyeva, Günel, et al.
Veröffentlicht: (2025)
MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Nuanced Emotion Recognition Based on a Segment-based MLLM Framework Leveraging Qwen3-Omni for AH Detection
von: Tang, Liang, et al.
Veröffentlicht: (2026)
von: Tang, Liang, et al.
Veröffentlicht: (2026)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
von: Li, Zuoou, et al.
Veröffentlicht: (2025)
von: Li, Zuoou, et al.
Veröffentlicht: (2025)
Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
von: Yan, Xudong, et al.
Veröffentlicht: (2024)
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
von: Du, Yifan, et al.
Veröffentlicht: (2025)
von: Du, Yifan, et al.
Veröffentlicht: (2025)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
START: Spatial and Textual Learning for Chart Understanding
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
BLOCK: An Open-Source Bi-Stage MLLM Character-to-Skin Pipeline for Minecraft
von: Guo, Hengquan
Veröffentlicht: (2026)
von: Guo, Hengquan
Veröffentlicht: (2026)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
von: Yu, An, et al.
Veröffentlicht: (2025)
von: Yu, An, et al.
Veröffentlicht: (2025)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
Task-conditioned Ensemble of Expert Models for Continuous Learning
von: Sharma, Renu, et al.
Veröffentlicht: (2025)
von: Sharma, Renu, et al.
Veröffentlicht: (2025)
Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
von: Chung, Minjae, et al.
Veröffentlicht: (2025)
von: Chung, Minjae, et al.
Veröffentlicht: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
von: Ding, Sihao, et al.
Veröffentlicht: (2025)
von: Ding, Sihao, et al.
Veröffentlicht: (2025)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
von: Sun, Yueru, et al.
Veröffentlicht: (2026)
von: Sun, Yueru, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition
von: Sony, Redwan, et al.
Veröffentlicht: (2025) -
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
von: Sony, Redwan, et al.
Veröffentlicht: (2025) -
A Parametric Approach to Adversarial Augmentation for Cross-Domain Iris Presentation Attack Detection
von: Pal, Debasmita, et al.
Veröffentlicht: (2024) -
Adversarial Watermarking for Face Recognition
von: Yao, Yuguang, et al.
Veröffentlicht: (2024) -
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026)