Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yiming, Yan, Qiangyu, Jiang, Borui, Han, Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914541649002496
author Zhang, Yiming
Yan, Qiangyu
Jiang, Borui
Han, Kai
author_facet Zhang, Yiming
Yan, Qiangyu
Jiang, Borui
Han, Kai
contents The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradigm, reliant on end-to-end generation or explicit, language-centric chains of thought (CoT), which can be inefficient, verbose, and prone to hallucination. This work posits that robust reasoning should evolve within a latent space, integrating multimodal signals seamlessly. We propose multimodal latent reasoning via HIerarchical Visual cuEs injection (\emph{HIVE}), a novel framework that instills deliberate, "slow thinking" without depending on superficial textual rationales. Our method recursively extends transformer blocks, creating an internal loop for iterative reasoning refinement. Crucially, it injectively grounds this process with hierarchical visual cues from global scene context to fine-grained regional details directly into the model's latent representations. This enables the model to perform grounded, multi-step inference entirely in the aligned latent space. Extensive evaluations demonstrate that test-time scaling is effective when incorporating vision knowledge, and that integrating hierarchical information significantly enhances the model's understanding of complex scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05359
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
Zhang, Yiming
Yan, Qiangyu
Jiang, Borui
Han, Kai
Computer Vision and Pattern Recognition
The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradigm, reliant on end-to-end generation or explicit, language-centric chains of thought (CoT), which can be inefficient, verbose, and prone to hallucination. This work posits that robust reasoning should evolve within a latent space, integrating multimodal signals seamlessly. We propose multimodal latent reasoning via HIerarchical Visual cuEs injection (\emph{HIVE}), a novel framework that instills deliberate, "slow thinking" without depending on superficial textual rationales. Our method recursively extends transformer blocks, creating an internal loop for iterative reasoning refinement. Crucially, it injectively grounds this process with hierarchical visual cues from global scene context to fine-grained regional details directly into the model's latent representations. This enables the model to perform grounded, multi-step inference entirely in the aligned latent space. Extensive evaluations demonstrate that test-time scaling is effective when incorporating vision knowledge, and that integrating hierarchical information significantly enhances the model's understanding of complex scenes.
title Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.05359