Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Chenfeng, He, Wei, Zhu, Xuhan, Zhou, Chunpeng, Li, Qizhen, Yan, Song, Zheng, Yufei, Yu, Chengjun, Lu, Fan, Zhai, Wei, Cao, Yang, Yu, Pengfei, Zha, Zheng-Jun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913122018656256
author Wang, Chenfeng
He, Wei
Zhu, Xuhan
Zhou, Chunpeng
Li, Qizhen
Yan, Song
Zheng, Yufei
Yu, Chengjun
Lu, Fan
Zhai, Wei
Cao, Yang
Yu, Pengfei
Zha, Zheng-Jun
author_facet Wang, Chenfeng
He, Wei
Zhu, Xuhan
Zhou, Chunpeng
Li, Qizhen
Yan, Song
Zheng, Yufei
Yu, Chengjun
Lu, Fan
Zhai, Wei
Cao, Yang
Yu, Pengfei
Zha, Zheng-Jun
contents In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent sequences. However, we discover a counterintuitive phenomenon: the performance of existing latent visual reasoning methods systematically degrades as the latent sequence grows longer. We reveal the root cause: Information Gain Collapse -- autoregressive generation makes each step highly dependent on prior outputs, so subsequent tokens can barely introduce new information. We further identify that heavily pooled ($\geq 128\times$) image embeddings used as supervision targets provide no more signal than meaningless placeholders. Motivated by these insights, we propose SCOLAR (Self-COnsistent LAtent Reasoning), which introduces a lightweight detransformer that leverages the LLM's full-sequence hidden states to generate auxiliary visual tokens in a single shot, with each token independently anchored to the original visual space. Combined with three-stage SFT and ALPO reinforcement learning, SCOLAR extends acceptable latent CoT length by over $30\times$, achieves state-of-the-art among open-source models on real-world reasoning benchmarks (+14.12% over backbone), and demonstrates strong out-of-distribution generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12163
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Wang, Chenfeng
He, Wei
Zhu, Xuhan
Zhou, Chunpeng
Li, Qizhen
Yan, Song
Zheng, Yufei
Yu, Chengjun
Lu, Fan
Zhai, Wei
Cao, Yang
Yu, Pengfei
Zha, Zheng-Jun
Computer Vision and Pattern Recognition
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent sequences. However, we discover a counterintuitive phenomenon: the performance of existing latent visual reasoning methods systematically degrades as the latent sequence grows longer. We reveal the root cause: Information Gain Collapse -- autoregressive generation makes each step highly dependent on prior outputs, so subsequent tokens can barely introduce new information. We further identify that heavily pooled ($\geq 128\times$) image embeddings used as supervision targets provide no more signal than meaningless placeholders. Motivated by these insights, we propose SCOLAR (Self-COnsistent LAtent Reasoning), which introduces a lightweight detransformer that leverages the LLM's full-sequence hidden states to generate auxiliary visual tokens in a single shot, with each token independently anchored to the original visual space. Combined with three-stage SFT and ALPO reinforcement learning, SCOLAR extends acceptable latent CoT length by over $30\times$, achieves state-of-the-art among open-source models on real-world reasoning benchmarks (+14.12% over backbone), and demonstrates strong out-of-distribution generalization.
title Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12163