Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Dengyang, Wang, Zanyi, Li, Hengzhuang, Dang, Sizhe, Ma, Teli, Wei, Wei, Dai, Guang, Zhang, Lei, Wang, Mengmeng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2504.15650
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909752053727232
author Jiang, Dengyang
Wang, Zanyi
Li, Hengzhuang
Dang, Sizhe
Ma, Teli
Wei, Wei
Dai, Guang
Zhang, Lei
Wang, Mengmeng
author_facet Jiang, Dengyang
Wang, Zanyi
Li, Hengzhuang
Dang, Sizhe
Ma, Teli
Wei, Wei
Dai, Guang
Zhang, Lei
Wang, Mengmeng
contents Building a generalized affordance grounding model to identify actionable regions on objects is vital for real-world applications. Existing methods to train the model can be divided into weakly and fully supervised ways. However, the former method requires a complex training framework design and can not infer new actions without an auxiliary prior. While the latter often struggle with limited annotated data and components trained from scratch despite being simpler. This study focuses on fully supervised affordance grounding and overcomes its limitations by proposing AffordanceSAM, which extends SAM's generalization capacity in segmentation to affordance grounding. Specifically, we design an affordance-adaption module and curate a coarse-to-fine annotated dataset called C2F-Aff to thoroughly transfer SAM's robust performance to affordance in a three-stage training manner. Experimental results confirm that AffordanceSAM achieves state-of-the-art (SOTA) performance on the AGD20K benchmark and exhibits strong generalized capacity.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AffordanceSAM: Segment Anything Once More in Affordance Grounding
Jiang, Dengyang
Wang, Zanyi
Li, Hengzhuang
Dang, Sizhe
Ma, Teli
Wei, Wei
Dai, Guang
Zhang, Lei
Wang, Mengmeng
Computer Vision and Pattern Recognition
Building a generalized affordance grounding model to identify actionable regions on objects is vital for real-world applications. Existing methods to train the model can be divided into weakly and fully supervised ways. However, the former method requires a complex training framework design and can not infer new actions without an auxiliary prior. While the latter often struggle with limited annotated data and components trained from scratch despite being simpler. This study focuses on fully supervised affordance grounding and overcomes its limitations by proposing AffordanceSAM, which extends SAM's generalization capacity in segmentation to affordance grounding. Specifically, we design an affordance-adaption module and curate a coarse-to-fine annotated dataset called C2F-Aff to thoroughly transfer SAM's robust performance to affordance in a three-stage training manner. Experimental results confirm that AffordanceSAM achieves state-of-the-art (SOTA) performance on the AGD20K benchmark and exhibits strong generalized capacity.
title AffordanceSAM: Segment Anything Once More in Affordance Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.15650