LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Linquan, Jiang, Tianxiang, Dong, Yifei, Yang, Haoyu, Zhang, Fengji, Meng, Shichaang, Xuan, Ai, Song, Linqi, Keung, Jacky
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908767421988864
author Wu, Linquan
Jiang, Tianxiang
Dong, Yifei
Yang, Haoyu
Zhang, Fengji
Meng, Shichaang
Xuan, Ai
Song, Linqi
Keung, Jacky
author_facet Wu, Linquan
Jiang, Tianxiang
Dong, Yifei
Yang, Haoyu
Zhang, Fengji
Meng, Shichaang
Xuan, Ai
Song, Linqi
Keung, Jacky
contents Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mimic a teacher's textual output while attending to fundamentally divergent visual regions, effectively relying on language priors rather than grounded perception. To bridge this, we propose LaViT, a framework that aligns latent visual thoughts rather than static embeddings. LaViT compels the student to autoregressively reconstruct the teacher's visual semantics and attention trajectories prior to text generation, employing a curriculum sensory gating mechanism to prevent shortcut learning. Extensive experiments show that LaViT significantly enhances visual grounding, achieving up to +16.9% gains on complex reasoning tasks and enabling a compact 3B model to outperform larger open-source variants and proprietary models like GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10129
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Wu, Linquan
Jiang, Tianxiang
Dong, Yifei
Yang, Haoyu
Zhang, Fengji
Meng, Shichaang
Xuan, Ai
Song, Linqi
Keung, Jacky
Computer Vision and Pattern Recognition
Artificial Intelligence
Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mimic a teacher's textual output while attending to fundamentally divergent visual regions, effectively relying on language priors rather than grounded perception. To bridge this, we propose LaViT, a framework that aligns latent visual thoughts rather than static embeddings. LaViT compels the student to autoregressively reconstruct the teacher's visual semantics and attention trajectories prior to text generation, employing a curriculum sensory gating mechanism to prevent shortcut learning. Extensive experiments show that LaViT significantly enhances visual grounding, achieving up to +16.9% gains on complex reasoning tasks and enabling a compact 3B model to outperform larger open-source variants and proprietary models like GPT-4o.
title LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.10129