Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908748925108224 |
|---|---|
| author | Wu, Mingyuan Li, Meitang Yang, Jingcheng Jiang, Jize Yan, Kaizhuo Li, Zhaoheng Yu, Hanchao Zhang, Minjia Nahrstedt, Klara |
| author_facet | Wu, Mingyuan Li, Meitang Yang, Jingcheng Jiang, Jize Yan, Kaizhuo Li, Zhaoheng Yu, Hanchao Zhang, Minjia Nahrstedt, Klara |
| contents | Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely attributed to emergent self correction and self verification behaviors often elicited through reinforcement learning (RL). In this work, we ask whether the same recipe transfers to vision language models (VLMs), especially RL finetuned variants that claim strong visual mathematical reasoning.
Through extensive evaluation, we reach three main findings that differ markedly from text only models. First, generation time capability matters more than verification and refinement: simple majority voting consistently and substantially outperforms verification centric strategies such as best of N with self verification. Second, behaviors often associated with RL tuned models at inference time, such as the 'Aha moment,' do not yield reliable reasoning performance improvements. Third, visual information is not effectively integrated into the model's self verification process.
Overall, our analysis highlights a key limitation: current RL trained VLMs derive limited benefit from self verification in the visual modality, which constrains the effectiveness of inference time scaling for visual mathematical reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_17417 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling? Wu, Mingyuan Li, Meitang Yang, Jingcheng Jiang, Jize Yan, Kaizhuo Li, Zhaoheng Yu, Hanchao Zhang, Minjia Nahrstedt, Klara Machine Learning Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely attributed to emergent self correction and self verification behaviors often elicited through reinforcement learning (RL). In this work, we ask whether the same recipe transfers to vision language models (VLMs), especially RL finetuned variants that claim strong visual mathematical reasoning. Through extensive evaluation, we reach three main findings that differ markedly from text only models. First, generation time capability matters more than verification and refinement: simple majority voting consistently and substantially outperforms verification centric strategies such as best of N with self verification. Second, behaviors often associated with RL tuned models at inference time, such as the 'Aha moment,' do not yield reliable reasoning performance improvements. Third, visual information is not effectively integrated into the model's self verification process. Overall, our analysis highlights a key limitation: current RL trained VLMs derive limited benefit from self verification in the visual modality, which constrains the effectiveness of inference time scaling for visual mathematical reasoning. |
| title | Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling? |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2506.17417 |