VGR: Visual Grounded Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jiacong, Kang, Zijian, Wang, Haochen, Jiang, Haiyong, Li, Jiawen, Wu, Bohong, Wang, Ya, Ran, Jiao, Liang, Xiao, Feng, Chao, Xiao, Jun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918476577243136
author Wang, Jiacong
Kang, Zijian
Wang, Haochen
Jiang, Haiyong
Li, Jiawen
Wu, Bohong
Wang, Ya
Ran, Jiao
Liang, Xiao
Feng, Chao
Xiao, Jun
author_facet Wang, Jiacong
Kang, Zijian
Wang, Haochen
Jiang, Haiyong
Li, Jiawen
Wu, Bohong
Wang, Ya
Ran, Jiao
Liang, Xiao
Feng, Chao
Xiao, Jun
contents In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or science domains. This narrow focus limits their ability to handle complex visual reasoning tasks that demand comprehensive understanding of image details. To address these limitations, this paper introduces VGR, a novel reasoning multimodal large language model (MLLM) with enhanced fine-grained visual perception capabilities. Unlike traditional MLLMs that answer the question or reasoning solely on the language space, our VGR first detects relevant regions that may help to solve problems, and then provides precise answers based on replayed image regions. To achieve this, we conduct a large-scale SFT dataset called VGR -SFT that contains reasoning data with mixed vision grounding and language deduction. The inference pipeline of VGR allows the model to choose bounding boxes for visual reference and a replay stage is introduced to integrates the corresponding regions into the reasoning process, enhancing multimodel comprehension. Experiments on the LLaVA-NeXT-7B baseline show that VGR achieves superior performance on multi-modal benchmarks requiring comprehensive image detail understanding. Compared to the baseline, VGR uses only 30\% of the image token count while delivering scores of +4.1 on MMStar, +7.1 on AI2D, and a +12.9 improvement on ChartQA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VGR: Visual Grounded Reasoning
Wang, Jiacong
Kang, Zijian
Wang, Haochen
Jiang, Haiyong
Li, Jiawen
Wu, Bohong
Wang, Ya
Ran, Jiao
Liang, Xiao
Feng, Chao
Xiao, Jun
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or science domains. This narrow focus limits their ability to handle complex visual reasoning tasks that demand comprehensive understanding of image details. To address these limitations, this paper introduces VGR, a novel reasoning multimodal large language model (MLLM) with enhanced fine-grained visual perception capabilities. Unlike traditional MLLMs that answer the question or reasoning solely on the language space, our VGR first detects relevant regions that may help to solve problems, and then provides precise answers based on replayed image regions. To achieve this, we conduct a large-scale SFT dataset called VGR -SFT that contains reasoning data with mixed vision grounding and language deduction. The inference pipeline of VGR allows the model to choose bounding boxes for visual reference and a replay stage is introduced to integrates the corresponding regions into the reasoning process, enhancing multimodel comprehension. Experiments on the LLaVA-NeXT-7B baseline show that VGR achieves superior performance on multi-modal benchmarks requiring comprehensive image detail understanding. Compared to the baseline, VGR uses only 30\% of the image token count while delivering scores of +4.1 on MMStar, +7.1 on AI2D, and a +12.9 improvement on ChartQA.
title VGR: Visual Grounded Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.11991