Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hanqing, Wang, Shaoyang, Zhong, Yiming, Yang, Zemin, Wang, Jiamin, Cui, Zhiqing, Yuan, Jiahao, Han, Yifan, Liu, Mingyu, Ma, Yuexin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914580940193792
author Wang, Hanqing
Wang, Shaoyang
Zhong, Yiming
Yang, Zemin
Wang, Jiamin
Cui, Zhiqing
Yuan, Jiahao
Han, Yifan
Liu, Mingyu
Ma, Yuexin
author_facet Wang, Hanqing
Wang, Shaoyang
Zhong, Yiming
Yang, Zemin
Wang, Jiamin
Cui, Zhiqing
Yuan, Jiahao
Han, Yifan
Liu, Mingyu
Ma, Yuexin
contents Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It plays a vital role in the fields of human-robot interaction, human-object interaction, embodied manipulation, and embodied perception. Existing models often neglect the affordance shared among different objects because they lack the Chain-of-Thought(CoT) reasoning abilities, limiting their out-of-domain (OOD) generalization and explicit reasoning capabilities. To address these challenges, we propose Affordance-R1, the first unified affordance grounding framework that integrates cognitive CoT guided Group Relative Policy Optimization (GRPO) within a reinforcement learning paradigm. Specifically, we designed a sophisticated affordance function, which contains format, perception, and cognition rewards to effectively guide optimization directions. Furthermore, we constructed a high-quality affordance-centric reasoning dataset, ReasonAff, to support training. Trained exclusively via reinforcement learning with GRPO and without explicit reasoning data, Affordance-R1 achieves robust zero-shot generalization and exhibits emergent test-time reasoning capabilities. Comprehensive experiments demonstrate that our model outperforms well-established methods and exhibits open-world generalization. To the best of our knowledge, Affordance-R1 is the first to integrate GRPO-based RL with reasoning into affordance reasoning. The code of our method and our dataset is released on https://github.com/hq-King/Affordance-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06206
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
Wang, Hanqing
Wang, Shaoyang
Zhong, Yiming
Yang, Zemin
Wang, Jiamin
Cui, Zhiqing
Yuan, Jiahao
Han, Yifan
Liu, Mingyu
Ma, Yuexin
Robotics
Computer Vision and Pattern Recognition
Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It plays a vital role in the fields of human-robot interaction, human-object interaction, embodied manipulation, and embodied perception. Existing models often neglect the affordance shared among different objects because they lack the Chain-of-Thought(CoT) reasoning abilities, limiting their out-of-domain (OOD) generalization and explicit reasoning capabilities. To address these challenges, we propose Affordance-R1, the first unified affordance grounding framework that integrates cognitive CoT guided Group Relative Policy Optimization (GRPO) within a reinforcement learning paradigm. Specifically, we designed a sophisticated affordance function, which contains format, perception, and cognition rewards to effectively guide optimization directions. Furthermore, we constructed a high-quality affordance-centric reasoning dataset, ReasonAff, to support training. Trained exclusively via reinforcement learning with GRPO and without explicit reasoning data, Affordance-R1 achieves robust zero-shot generalization and exhibits emergent test-time reasoning capabilities. Comprehensive experiments demonstrate that our model outperforms well-established methods and exhibits open-world generalization. To the best of our knowledge, Affordance-R1 is the first to integrate GRPO-based RL with reasoning into affordance reasoning. The code of our method and our dataset is released on https://github.com/hq-King/Affordance-R1.
title Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06206