VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Divake, Tayebati, Sina, Naik, Devashri, Krishnan, Ranganath, Trivedi, Amit Ranjan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
by: Kumar, Divake, et al.
Published: (2025)
by: Kumar, Divake, et al.
Published: (2025)
EigenShield: Causal Subspace Filtering via Random Matrix Theory for Adversarially Robust Vision-Language Models
by: Darabi, Nastaran, et al.
Published: (2025)
by: Darabi, Nastaran, et al.
Published: (2025)
INTACT: Inducing Noise Tolerance through Adversarial Curriculum Training for LiDAR-based Safety-Critical Perception and Autonomy
by: Darabi, Nastaran, et al.
Published: (2025)
by: Darabi, Nastaran, et al.
Published: (2025)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
by: Jayasuriya, Dinithi, et al.
Published: (2025)
by: Jayasuriya, Dinithi, et al.
Published: (2025)
Learnable Conformal Prediction with Context-Aware Nonconformity Functions for Robotic Planning and Perception
by: Kumar, Divake, et al.
Published: (2025)
by: Kumar, Divake, et al.
Published: (2025)
TRIAGE: Type-Routed Interventions via Aleatoric-Epistemic Gated Estimation in Robotic Manipulation and Adaptive Perception -- Don't Treat All Uncertainty the Same
by: Kumar, Divake, et al.
Published: (2026)
by: Kumar, Divake, et al.
Published: (2026)
Uncertainty-Guided Inference-Time Depth Adaptation for Transformer-Based Visual Tracking
by: Poggi, Patrick, et al.
Published: (2026)
by: Poggi, Patrick, et al.
Published: (2026)
Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models
by: Tayebati, Sina, et al.
Published: (2025)
by: Tayebati, Sina, et al.
Published: (2025)
TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
by: Tayebati, Sina, et al.
Published: (2026)
by: Tayebati, Sina, et al.
Published: (2026)
Causal-Guided Dimension Reduction for Efficient Pareto Optimization
by: Jayasuriya, Dinithi, et al.
Published: (2025)
by: Jayasuriya, Dinithi, et al.
Published: (2025)
Belief Dynamics for Detecting Behavioral Shifts in Safe Collaborative Manipulation
by: Naik, Devashri, et al.
Published: (2026)
by: Naik, Devashri, et al.
Published: (2026)
Sense Less, Generate More: Pre-training LiDAR Perception with Masked Autoencoders for Ultra-Efficient 3D Sensing
by: Tayebati, Sina, et al.
Published: (2024)
by: Tayebati, Sina, et al.
Published: (2024)
Beyond Confidence: Adaptive Abstention in Dual-Threshold Conformal Prediction for Autonomous System Perception
by: Kumar, Divake, et al.
Published: (2025)
by: Kumar, Divake, et al.
Published: (2025)
Enhancing 3D Robotic Vision Robustness by Minimizing Adversarial Mutual Information through a Curriculum Training Approach
by: Darabi, Nastaran, et al.
Published: (2024)
by: Darabi, Nastaran, et al.
Published: (2024)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
Intelligent Sensing-to-Action for Robust Autonomy at the Edge: Opportunities and Challenges
by: Trivedi, Amit Ranjan, et al.
Published: (2025)
by: Trivedi, Amit Ranjan, et al.
Published: (2025)
M-PACE: Mother Child Framework for Multimodal Compliance
by: Verma, Shreyash, et al.
Published: (2025)
by: Verma, Shreyash, et al.
Published: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
by: Khangaonkar, Om, et al.
Published: (2026)
by: Khangaonkar, Om, et al.
Published: (2026)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
Re:Verse -- Can Your VLM Read a Manga?
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
by: Krishnan, Ranganath, et al.
Published: (2024)
by: Krishnan, Ranganath, et al.
Published: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration
by: Narayanan, Athmanarayanan Lakshmi, et al.
Published: (2025)
by: Narayanan, Athmanarayanan Lakshmi, et al.
Published: (2025)
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs
by: Ettori, Davide, et al.
Published: (2025)
by: Ettori, Davide, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
by: Meng, Rui, et al.
Published: (2025)
by: Meng, Rui, et al.
Published: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
by: Zhao, Pu, et al.
Published: (2025)
by: Zhao, Pu, et al.
Published: (2025)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
VLM-KG: Multimodal Radiology Knowledge Graph Generation
by: Abdullah, Abdullah, et al.
Published: (2025)
by: Abdullah, Abdullah, et al.
Published: (2025)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Similar Items
-
Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
by: Kumar, Divake, et al.
Published: (2025) -
EigenShield: Causal Subspace Filtering via Random Matrix Theory for Adversarially Robust Vision-Language Models
by: Darabi, Nastaran, et al.
Published: (2025) -
INTACT: Inducing Noise Tolerance through Adversarial Curriculum Training for LiDAR-based Safety-Critical Perception and Autonomy
by: Darabi, Nastaran, et al.
Published: (2025) -
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
by: Jayasuriya, Dinithi, et al.
Published: (2025) -
Learnable Conformal Prediction with Context-Aware Nonconformity Functions for Robotic Planning and Perception
by: Kumar, Divake, et al.
Published: (2025)