Interleaved Latent Visual Reasoning with Selective Perceptual Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Shuai, Wang, Siyuan, Liu, Xingyu, Li, Chenglin, Hou, Haowen, Wei, Zhongyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918297699614720
author Dong, Shuai
Wang, Siyuan
Liu, Xingyu
Li, Chenglin
Hou, Haowen
Wei, Zhongyu
author_facet Dong, Shuai
Wang, Siyuan
Liu, Xingyu
Li, Chenglin
Hou, Haowen
Wei, Zhongyu
contents Interleaved reasoning paradigms enhance Multimodal Large Language Models (MLLMs) with visual feedback but are hindered by the prohibitive computational cost of re-encoding pixel-dense images. A promising alternative, latent visual reasoning, circumvents this bottleneck yet faces limitations: methods either fail to capture intermediate state evolution due to single-step, non-interleaved structures, or sacrifice precise perceptual modeling by over-compressing features. We introduce Interleaved Latent Visual Reasoning (ILVR), a framework that unifies dynamic state evolution with precise perceptual modeling. ILVR interleaves textual generation with latent visual representations that act as specific, evolving cues for subsequent reasoning. Specifically, we employ a self-supervision strategy where a momentum teacher model selectively distills relevant features from ground-truth intermediate images into sparse supervision targets. This adaptive selection mechanism guides the model to autonomously generate context-aware visual signals. Extensive experiments on multimodal reasoning benchmarks demonstrate that ILVR outperforms existing approaches, effectively bridging the gap between fine-grained perception and sequential multimodal reasoning. The code is available at https://github.com/XD111ds/ILVR.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
Dong, Shuai
Wang, Siyuan
Liu, Xingyu
Li, Chenglin
Hou, Haowen
Wei, Zhongyu
Computation and Language
Computer Vision and Pattern Recognition
Interleaved reasoning paradigms enhance Multimodal Large Language Models (MLLMs) with visual feedback but are hindered by the prohibitive computational cost of re-encoding pixel-dense images. A promising alternative, latent visual reasoning, circumvents this bottleneck yet faces limitations: methods either fail to capture intermediate state evolution due to single-step, non-interleaved structures, or sacrifice precise perceptual modeling by over-compressing features. We introduce Interleaved Latent Visual Reasoning (ILVR), a framework that unifies dynamic state evolution with precise perceptual modeling. ILVR interleaves textual generation with latent visual representations that act as specific, evolving cues for subsequent reasoning. Specifically, we employ a self-supervision strategy where a momentum teacher model selectively distills relevant features from ground-truth intermediate images into sparse supervision targets. This adaptive selection mechanism guides the model to autonomously generate context-aware visual signals. Extensive experiments on multimodal reasoning benchmarks demonstrate that ILVR outperforms existing approaches, effectively bridging the gap between fine-grained perception and sequential multimodal reasoning. The code is available at https://github.com/XD111ds/ILVR.
title Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05665