LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908767421988864 |
|---|---|
| author | Wu, Linquan Jiang, Tianxiang Dong, Yifei Yang, Haoyu Zhang, Fengji Meng, Shichaang Xuan, Ai Song, Linqi Keung, Jacky |
| author_facet | Wu, Linquan Jiang, Tianxiang Dong, Yifei Yang, Haoyu Zhang, Fengji Meng, Shichaang Xuan, Ai Song, Linqi Keung, Jacky |
| contents | Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mimic a teacher's textual output while attending to fundamentally divergent visual regions, effectively relying on language priors rather than grounded perception. To bridge this, we propose LaViT, a framework that aligns latent visual thoughts rather than static embeddings. LaViT compels the student to autoregressively reconstruct the teacher's visual semantics and attention trajectories prior to text generation, employing a curriculum sensory gating mechanism to prevent shortcut learning. Extensive experiments show that LaViT significantly enhances visual grounding, achieving up to +16.9% gains on complex reasoning tasks and enabling a compact 3B model to outperform larger open-source variants and proprietary models like GPT-4o. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_10129 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning Wu, Linquan Jiang, Tianxiang Dong, Yifei Yang, Haoyu Zhang, Fengji Meng, Shichaang Xuan, Ai Song, Linqi Keung, Jacky Computer Vision and Pattern Recognition Artificial Intelligence Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mimic a teacher's textual output while attending to fundamentally divergent visual regions, effectively relying on language priors rather than grounded perception. To bridge this, we propose LaViT, a framework that aligns latent visual thoughts rather than static embeddings. LaViT compels the student to autoregressively reconstruct the teacher's visual semantics and attention trajectories prior to text generation, employing a curriculum sensory gating mechanism to prevent shortcut learning. Extensive experiments show that LaViT significantly enhances visual grounding, achieving up to +16.9% gains on complex reasoning tasks and enabling a compact 3B model to outperform larger open-source variants and proprietary models like GPT-4o. |
| title | LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2601.10129 |