CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Ziyang, Meng, Linjian, Wu, Yiming, Li, Yuhan, Liu, Yuhao, Zhao, Zhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916003182542848
author Ding, Ziyang
Meng, Linjian
Wu, Yiming
Li, Yuhan
Liu, Yuhao
Zhao, Zhen
author_facet Ding, Ziyang
Meng, Linjian
Wu, Yiming
Li, Yuhan
Liu, Yuhao
Zhao, Zhen
contents Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle-based perturbation, which expands the semantic latent space and avoids over-constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post-training to enable fine-grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar. The data, codes, and models are released at https://github.com/Oscar-dzy/CoLVR.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08802
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization
Ding, Ziyang
Meng, Linjian
Wu, Yiming
Li, Yuhan
Liu, Yuhao
Zhao, Zhen
Computer Vision and Pattern Recognition
Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle-based perturbation, which expands the semantic latent space and avoids over-constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post-training to enable fine-grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar. The data, codes, and models are released at https://github.com/Oscar-dzy/CoLVR.
title CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.08802