v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911653565562880 |
|---|---|
| author | Chung, Jiwan Kim, Junhyeok Kim, Siyeol Lee, Jaeyoung Kim, Min Soo Yu, Youngjae |
| author_facet | Chung, Jiwan Kim, Junhyeok Kim, Siyeol Lee, Jaeyoung Kim, Min Soo Yu, Youngjae |
| contents | When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to key-value cache and then reason purely in text, making it hard to re-ground intermediate steps. We empirically confirm this: as reasoning chains lengthen, models progressively lose focus on relevant regions. We introduce v1, a lightweight extension for active visual referencing via point-and-copy: the model selects relevant image patches and copies their embeddings back into the reasoning stream. Crucially, our point-and-copy mechanism retrieves patches using their semantic representations as keys, ensuring perceptual evidence remains aligned with the reasoning space. To train this behavior, we build v1g, a dataset of 300K multimodal reasoning traces with interleaved grounding annotations. Across multimodal mathematical reasoning benchmarks, v1 consistently outperforms comparable baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18842 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning Chung, Jiwan Kim, Junhyeok Kim, Siyeol Lee, Jaeyoung Kim, Min Soo Yu, Youngjae Computation and Language Computer Vision and Pattern Recognition When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to key-value cache and then reason purely in text, making it hard to re-ground intermediate steps. We empirically confirm this: as reasoning chains lengthen, models progressively lose focus on relevant regions. We introduce v1, a lightweight extension for active visual referencing via point-and-copy: the model selects relevant image patches and copies their embeddings back into the reasoning stream. Crucially, our point-and-copy mechanism retrieves patches using their semantic representations as keys, ensuring perceptual evidence remains aligned with the reasoning space. To train this behavior, we build v1g, a dataset of 300K multimodal reasoning traces with interleaved grounding annotations. Across multimodal mathematical reasoning benchmarks, v1 consistently outperforms comparable baselines. |
| title | v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.18842 |