Learning Self-Correction in Vision-Language Models via Rollout Augmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ding, Yi, Qiu, Ziliang, Li, Bolian, Zhang, Ruqi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914315617959936
author Ding, Yi
Qiu, Ziliang
Li, Bolian
Zhang, Ruqi
author_facet Ding, Yi
Qiu, Ziliang
Li, Bolian
Zhang, Ruqi
contents Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning (RL) methods struggle to learn it, as effective self-correction behaviors emerge only rarely, making learning signals extremely sparse. To address this challenge, we propose correction-specific rollouts (Octopus), an RL rollout augmentation framework that synthesizes dense self-correction examples by recombining existing rollouts. This augmentation simultaneously improves sample efficiency due to rollout reuse and stabilizes RL optimization through balanced supervision. Furthermore, we introduce a response-masking strategy that decouples self-correction from direct reasoning, avoiding signal conflicts and enabling both behaviors to be learned effectively. Building on this, we introduce Octopus-8B, a reasoning VLM with controllable self-correction capability. Across 7 benchmarks, it achieves SoTA performance among open-source VLMs, outperforming the best RLVR baseline by 1.0 score while requiring only $0.72\times$ training time per step.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08503
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Self-Correction in Vision-Language Models via Rollout Augmentation
Ding, Yi
Qiu, Ziliang
Li, Bolian
Zhang, Ruqi
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning (RL) methods struggle to learn it, as effective self-correction behaviors emerge only rarely, making learning signals extremely sparse. To address this challenge, we propose correction-specific rollouts (Octopus), an RL rollout augmentation framework that synthesizes dense self-correction examples by recombining existing rollouts. This augmentation simultaneously improves sample efficiency due to rollout reuse and stabilizes RL optimization through balanced supervision. Furthermore, we introduce a response-masking strategy that decouples self-correction from direct reasoning, avoiding signal conflicts and enabling both behaviors to be learned effectively. Building on this, we introduce Octopus-8B, a reasoning VLM with controllable self-correction capability. Across 7 benchmarks, it achieves SoTA performance among open-source VLMs, outperforming the best RLVR baseline by 1.0 score while requiring only $0.72\times$ training time per step.
title Learning Self-Correction in Vision-Language Models via Rollout Augmentation
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.08503