History-Guided Iterative Visual Reasoning with Self-Correction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Xinglong, Peng, Zhilin, Liu, Zhanzhan, Shi, Haochen, Huang, Sheng-Jun
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917247523487744
author Yang, Xinglong
Peng, Zhilin
Liu, Zhanzhan
Shi, Haochen
Huang, Sheng-Jun
author_facet Yang, Xinglong
Peng, Zhilin
Liu, Zhanzhan
Shi, Haochen
Huang, Sheng-Jun
contents Self-consistency methods are the core technique for improving the reasoning reliability of multimodal large language models (MLLMs). By generating multiple reasoning results through repeated sampling and selecting the best answer via voting, they play an important role in cross-modal tasks. However, most existing self-consistency methods are limited to a fixed ``repeated sampling and voting'' paradigm and do not reuse historical reasoning information. As a result, models struggle to actively correct visual understanding errors and dynamically adjust their reasoning during iteration. Inspired by the human reasoning behavior of repeated verification and dynamic error correction, we propose the H-GIVR framework. During iterative reasoning, the MLLM observes the image multiple times and uses previously generated answers as references for subsequent steps, enabling dynamic correction of errors and improving answer accuracy. We conduct comprehensive experiments on five datasets and three models. The results show that the H-GIVR framework can significantly improve cross-modal reasoning accuracy while maintaining low computational cost. For instance, using \texttt{Llama3.2-vision:11b} on the ScienceQA dataset, the model requires an average of 2.57 responses per question to achieve an accuracy of 78.90\%, representing a 107\% improvement over the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04413
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle History-Guided Iterative Visual Reasoning with Self-Correction
Yang, Xinglong
Peng, Zhilin
Liu, Zhanzhan
Shi, Haochen
Huang, Sheng-Jun
Computation and Language
Artificial Intelligence
Multimedia
Self-consistency methods are the core technique for improving the reasoning reliability of multimodal large language models (MLLMs). By generating multiple reasoning results through repeated sampling and selecting the best answer via voting, they play an important role in cross-modal tasks. However, most existing self-consistency methods are limited to a fixed ``repeated sampling and voting'' paradigm and do not reuse historical reasoning information. As a result, models struggle to actively correct visual understanding errors and dynamically adjust their reasoning during iteration. Inspired by the human reasoning behavior of repeated verification and dynamic error correction, we propose the H-GIVR framework. During iterative reasoning, the MLLM observes the image multiple times and uses previously generated answers as references for subsequent steps, enabling dynamic correction of errors and improving answer accuracy. We conduct comprehensive experiments on five datasets and three models. The results show that the H-GIVR framework can significantly improve cross-modal reasoning accuracy while maintaining low computational cost. For instance, using \texttt{Llama3.2-vision:11b} on the ScienceQA dataset, the model requires an average of 2.57 responses per question to achieve an accuracy of 78.90\%, representing a 107\% improvement over the baseline.
title History-Guided Iterative Visual Reasoning with Self-Correction
topic Computation and Language
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2602.04413