Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Mingyuan, Li, Meitang, Yang, Jingcheng, Jiang, Jize, Yan, Kaizhuo, Li, Zhaoheng, Yu, Hanchao, Zhang, Minjia, Nahrstedt, Klara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908748925108224
author Wu, Mingyuan
Li, Meitang
Yang, Jingcheng
Jiang, Jize
Yan, Kaizhuo
Li, Zhaoheng
Yu, Hanchao
Zhang, Minjia
Nahrstedt, Klara
author_facet Wu, Mingyuan
Li, Meitang
Yang, Jingcheng
Jiang, Jize
Yan, Kaizhuo
Li, Zhaoheng
Yu, Hanchao
Zhang, Minjia
Nahrstedt, Klara
contents Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely attributed to emergent self correction and self verification behaviors often elicited through reinforcement learning (RL). In this work, we ask whether the same recipe transfers to vision language models (VLMs), especially RL finetuned variants that claim strong visual mathematical reasoning. Through extensive evaluation, we reach three main findings that differ markedly from text only models. First, generation time capability matters more than verification and refinement: simple majority voting consistently and substantially outperforms verification centric strategies such as best of N with self verification. Second, behaviors often associated with RL tuned models at inference time, such as the 'Aha moment,' do not yield reliable reasoning performance improvements. Third, visual information is not effectively integrated into the model's self verification process. Overall, our analysis highlights a key limitation: current RL trained VLMs derive limited benefit from self verification in the visual modality, which constrains the effectiveness of inference time scaling for visual mathematical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Wu, Mingyuan
Li, Meitang
Yang, Jingcheng
Jiang, Jize
Yan, Kaizhuo
Li, Zhaoheng
Yu, Hanchao
Zhang, Minjia
Nahrstedt, Klara
Machine Learning
Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely attributed to emergent self correction and self verification behaviors often elicited through reinforcement learning (RL). In this work, we ask whether the same recipe transfers to vision language models (VLMs), especially RL finetuned variants that claim strong visual mathematical reasoning. Through extensive evaluation, we reach three main findings that differ markedly from text only models. First, generation time capability matters more than verification and refinement: simple majority voting consistently and substantially outperforms verification centric strategies such as best of N with self verification. Second, behaviors often associated with RL tuned models at inference time, such as the 'Aha moment,' do not yield reliable reasoning performance improvements. Third, visual information is not effectively integrated into the model's self verification process. Overall, our analysis highlights a key limitation: current RL trained VLMs derive limited benefit from self verification in the visual modality, which constrains the effectiveness of inference time scaling for visual mathematical reasoning.
title Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
topic Machine Learning
url https://arxiv.org/abs/2506.17417