v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chung, Jiwan, Kim, Junhyeok, Kim, Siyeol, Lee, Jaeyoung, Kim, Min Soo, Yu, Youngjae
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911653565562880
author Chung, Jiwan
Kim, Junhyeok
Kim, Siyeol
Lee, Jaeyoung
Kim, Min Soo
Yu, Youngjae
author_facet Chung, Jiwan
Kim, Junhyeok
Kim, Siyeol
Lee, Jaeyoung
Kim, Min Soo
Yu, Youngjae
contents When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to key-value cache and then reason purely in text, making it hard to re-ground intermediate steps. We empirically confirm this: as reasoning chains lengthen, models progressively lose focus on relevant regions. We introduce v1, a lightweight extension for active visual referencing via point-and-copy: the model selects relevant image patches and copies their embeddings back into the reasoning stream. Crucially, our point-and-copy mechanism retrieves patches using their semantic representations as keys, ensuring perceptual evidence remains aligned with the reasoning space. To train this behavior, we build v1g, a dataset of 300K multimodal reasoning traces with interleaved grounding annotations. Across multimodal mathematical reasoning benchmarks, v1 consistently outperforms comparable baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
Chung, Jiwan
Kim, Junhyeok
Kim, Siyeol
Lee, Jaeyoung
Kim, Min Soo
Yu, Youngjae
Computation and Language
Computer Vision and Pattern Recognition
When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to key-value cache and then reason purely in text, making it hard to re-ground intermediate steps. We empirically confirm this: as reasoning chains lengthen, models progressively lose focus on relevant regions. We introduce v1, a lightweight extension for active visual referencing via point-and-copy: the model selects relevant image patches and copies their embeddings back into the reasoning stream. Crucially, our point-and-copy mechanism retrieves patches using their semantic representations as keys, ensuring perceptual evidence remains aligned with the reasoning space. To train this behavior, we build v1g, a dataset of 300K multimodal reasoning traces with interleaved grounding annotations. Across multimodal mathematical reasoning benchmarks, v1 consistently outperforms comparable baselines.
title v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18842