From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuyuan, Ji, Yiping, Le, Anjie, Zhu, Jiayuan, Pan, Jiazhen, Peng, Can, Deng, Jiajun, Liu, Fengbei, Wu, Junde
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917500151660544
author Liu, Yuyuan
Ji, Yiping
Le, Anjie
Zhu, Jiayuan
Pan, Jiazhen
Peng, Can
Deng, Jiajun
Liu, Fengbei
Wu, Junde
author_facet Liu, Yuyuan
Ji, Yiping
Le, Anjie
Zhu, Jiayuan
Pan, Jiazhen
Peng, Can
Deng, Jiajun
Liu, Fengbei
Wu, Junde
contents Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing methods, mainly based on GRPO, assign rewards at the response level. Such sparse reward, often criterion-induced, leads to minimal learning signals when all candidate responses fail in challenging scenarios. In this work, we propose a group-revision optimisation paradigm that enhances learning on hard cases. It begins with a sampled initial response and generates a set of revised candidates to explore improved grounding outcomes. Inspired by reward shaping, we introduce a consolidation process that quantifies each candidate's improvement over the initial attempt and converts it into informative shaping signals. These signals are used to both refine the reward and modulate the advantage, amplifying the influence of high-quality revisions. Our method achieves consistent gains across referring and reasoning segmentation, REC, and counting benchmarks compared with prior GRPO-based models. Our code is available at https://github.com/yyliu01/GroupRevision.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15951
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding
Liu, Yuyuan
Ji, Yiping
Le, Anjie
Zhu, Jiayuan
Pan, Jiazhen
Peng, Can
Deng, Jiajun
Liu, Fengbei
Wu, Junde
Computer Vision and Pattern Recognition
Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing methods, mainly based on GRPO, assign rewards at the response level. Such sparse reward, often criterion-induced, leads to minimal learning signals when all candidate responses fail in challenging scenarios. In this work, we propose a group-revision optimisation paradigm that enhances learning on hard cases. It begins with a sampled initial response and generates a set of revised candidates to explore improved grounding outcomes. Inspired by reward shaping, we introduce a consolidation process that quantifies each candidate's improvement over the initial attempt and converts it into informative shaping signals. These signals are used to both refine the reward and modulate the advantage, amplifying the influence of high-quality revisions. Our method achieves consistent gains across referring and reasoning segmentation, REC, and counting benchmarks compared with prior GRPO-based models. Our code is available at https://github.com/yyliu01/GroupRevision.
title From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.15951