Linking Perception, Confidence and Accuracy in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Yuetian, Wang, Yucheng, Zhang, Rongyu, Xu, Zhijie, Yang, Boyu, Kong, Ming, Liu, Jie, Zhu, Qiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917335716069376
author Du, Yuetian
Wang, Yucheng
Zhang, Rongyu
Xu, Zhijie
Yang, Boyu
Kong, Ming
Liu, Jie
Zhu, Qiang
author_facet Du, Yuetian
Wang, Yucheng
Zhang, Rongyu
Xu, Zhijie
Yang, Boyu
Kong, Ming
Liu, Jie
Zhu, Qiang
contents Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual perception to improve accuracy. However, a critical question remains unexplored: Do models know when they do not know? Through a probing experiment, we reveal a severe confidence miscalibration problem in MLLMs. To address this, we propose Confidence-Driven Reinforcement Learning (CDRL), which uses original-noise image pairs and a novel confidence-based reward to enhance perceptual sensitivity and robustly calibrate the model's confidence. Beyond training benefits, calibrated confidence enables more effective test-time scaling as a free lunch. We further propose Confidence-Aware Test-Time Scaling (CA-TTS), which dynamically coordinates Self-Consistency, Self-Reflection, and Visual Self-Check modules guided by confidence signals. An Expert Model acts in multiple roles (e.g., Planner, Critic, Voter) to schedule these modules and provide external verification. Our integrated framework establishes new state-of-the-art results with consistent 8.8% gains across four benchmarks. More ablation studies demonstrate the effectiveness of each module and scaling superiority.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12149
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Linking Perception, Confidence and Accuracy in MLLMs
Du, Yuetian
Wang, Yucheng
Zhang, Rongyu
Xu, Zhijie
Yang, Boyu
Kong, Ming
Liu, Jie
Zhu, Qiang
Computer Vision and Pattern Recognition
Computation and Language
Recent advances in Multi-modal Large Language Models (MLLMs) have predominantly focused on enhancing visual perception to improve accuracy. However, a critical question remains unexplored: Do models know when they do not know? Through a probing experiment, we reveal a severe confidence miscalibration problem in MLLMs. To address this, we propose Confidence-Driven Reinforcement Learning (CDRL), which uses original-noise image pairs and a novel confidence-based reward to enhance perceptual sensitivity and robustly calibrate the model's confidence. Beyond training benefits, calibrated confidence enables more effective test-time scaling as a free lunch. We further propose Confidence-Aware Test-Time Scaling (CA-TTS), which dynamically coordinates Self-Consistency, Self-Reflection, and Visual Self-Check modules guided by confidence signals. An Expert Model acts in multiple roles (e.g., Planner, Critic, Voter) to schedule these modules and provide external verification. Our integrated framework establishes new state-of-the-art results with consistent 8.8% gains across four benchmarks. More ablation studies demonstrate the effectiveness of each module and scaling superiority.
title Linking Perception, Confidence and Accuracy in MLLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2603.12149