Latent Implicit Visual Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Kelvin, Shang, Chuyi, Karlinsky, Leonid, Feris, Rogerio, Darrell, Trevor, Herzig, Roei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915693455212544
author Li, Kelvin
Shang, Chuyi
Karlinsky, Leonid
Feris, Rogerio
Darrell, Trevor
Herzig, Roei
author_facet Li, Kelvin
Shang, Chuyi
Karlinsky, Leonid
Feris, Rogerio
Darrell, Trevor
Herzig, Roei
contents While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are limited in their ability to handle reasoning tasks that are predominantly visual. Recent approaches have sought to address this by supervising intermediate visual steps with helper images, depth maps, or image crops. However, these strategies impose restrictive priors on what "useful" visual abstractions look like, add heavy annotation costs, and struggle to generalize across tasks. To address this critical limitation, we propose a task-agnostic mechanism that trains LMMs to discover and use visual reasoning tokens without explicit supervision. These tokens attend globally and re-encode the image in a task-adaptive way, enabling the model to extract relevant visual information without hand-crafted supervision. Our approach outperforms direct fine-tuning and achieves state-of-the-art results on a diverse range of vision-centric tasks -- including those where intermediate abstractions are hard to specify -- while also generalizing to multi-task instruction tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Implicit Visual Reasoning
Li, Kelvin
Shang, Chuyi
Karlinsky, Leonid
Feris, Rogerio
Darrell, Trevor
Herzig, Roei
Computer Vision and Pattern Recognition
While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning modality. As a result, they are limited in their ability to handle reasoning tasks that are predominantly visual. Recent approaches have sought to address this by supervising intermediate visual steps with helper images, depth maps, or image crops. However, these strategies impose restrictive priors on what "useful" visual abstractions look like, add heavy annotation costs, and struggle to generalize across tasks. To address this critical limitation, we propose a task-agnostic mechanism that trains LMMs to discover and use visual reasoning tokens without explicit supervision. These tokens attend globally and re-encode the image in a task-adaptive way, enabling the model to extract relevant visual information without hand-crafted supervision. Our approach outperforms direct fine-tuning and achieves state-of-the-art results on a diverse range of vision-centric tasks -- including those where intermediate abstractions are hard to specify -- while also generalizing to multi-task instruction tuning.
title Latent Implicit Visual Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21218