DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hanqing, Zhang, Zhenhao, Ji, Kaiyang, Liu, Mingyu, Yin, Wenti, Chen, Yuchao, Liu, Zhirui, Zeng, Xiangyu, Gui, Tianxiang, Zhang, Hangxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909719244832768
author Wang, Hanqing
Zhang, Zhenhao
Ji, Kaiyang
Liu, Mingyu
Yin, Wenti
Chen, Yuchao
Liu, Zhirui
Zeng, Xiangyu
Gui, Tianxiang
Zhang, Hangxing
author_facet Wang, Hanqing
Zhang, Zhenhao
Ji, Kaiyang
Liu, Mingyu
Yin, Wenti
Chen, Yuchao
Liu, Zhirui
Zeng, Xiangyu
Gui, Tianxiang
Zhang, Hangxing
contents 3D object affordance grounding aims to predict the touchable regions on a 3d object, which is crucial for human-object interaction, human-robot interaction, embodied perception, and robot learning. Recent advances tackle this problem via learning from demonstration images. However, these methods fail to capture the general affordance knowledge within the image, leading to poor generalization. To address this issue, we propose to use text-to-image diffusion models to extract the general affordance knowledge because we find that such models can generate semantically valid HOI images, which demonstrate that their internal representation space is highly correlated with real-world affordance concepts. Specifically, we introduce the DAG, a diffusion-based 3d affordance grounding framework, which leverages the frozen internal representations of the text-to-image diffusion model and unlocks affordance knowledge within the diffusion model to perform 3D affordance grounding. We further introduce an affordance block and a multi-source affordance decoder to endow 3D dense affordance prediction. Extensive experimental evaluations show that our model excels over well-established methods and exhibits open-world generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01651
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
Wang, Hanqing
Zhang, Zhenhao
Ji, Kaiyang
Liu, Mingyu
Yin, Wenti
Chen, Yuchao
Liu, Zhirui
Zeng, Xiangyu
Gui, Tianxiang
Zhang, Hangxing
Computer Vision and Pattern Recognition
3D object affordance grounding aims to predict the touchable regions on a 3d object, which is crucial for human-object interaction, human-robot interaction, embodied perception, and robot learning. Recent advances tackle this problem via learning from demonstration images. However, these methods fail to capture the general affordance knowledge within the image, leading to poor generalization. To address this issue, we propose to use text-to-image diffusion models to extract the general affordance knowledge because we find that such models can generate semantically valid HOI images, which demonstrate that their internal representation space is highly correlated with real-world affordance concepts. Specifically, we introduce the DAG, a diffusion-based 3d affordance grounding framework, which leverages the frozen internal representations of the text-to-image diffusion model and unlocks affordance knowledge within the diffusion model to perform 3D affordance grounding. We further introduce an affordance block and a multi-source affordance decoder to endow 3D dense affordance prediction. Extensive experimental evaluations show that our model excels over well-established methods and exhibits open-world generalization.
title DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.01651