DAG: Unleash the Potential of Diffusion Model for Open-Vocabulary 3D Affordance Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hanqing, Zhang, Zhenhao, Ji, Kaiyang, Liu, Mingyu, Yin, Wenti, Chen, Yuchao, Liu, Zhirui, Zeng, Xiangyu, Gui, Tianxiang, Zhang, Hangxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026)
by: Wang, Hanqing, et al.
Published: (2026)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
by: Tong, Edmond, et al.
Published: (2024)
by: Tong, Edmond, et al.
Published: (2024)
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024)
by: Shao, Yawen, et al.
Published: (2024)
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
by: Liu, Zhirui, et al.
Published: (2025)
by: Liu, Zhirui, et al.
Published: (2025)
FSAG: Enhancing Human-to-Dexterous-Hand Finger-Specific Affordance Grounding via Diffusion Models
by: Han, Yifan, et al.
Published: (2026)
by: Han, Yifan, et al.
Published: (2026)
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
by: Lee, Junha, et al.
Published: (2025)
by: Lee, Junha, et al.
Published: (2025)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment
by: Gou, Dongqiang, et al.
Published: (2026)
by: Gou, Dongqiang, et al.
Published: (2026)
OVAL-Grasp: Open-Vocabulary Affordance Localization for Task Oriented Grasping
by: Tong, Edmond, et al.
Published: (2025)
by: Tong, Edmond, et al.
Published: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
by: Sun, Haowen, et al.
Published: (2026)
by: Sun, Haowen, et al.
Published: (2026)
Open-Vocabulary Calibration for Fine-tuned CLIP
by: Wang, Shuoyuan, et al.
Published: (2024)
by: Wang, Shuoyuan, et al.
Published: (2024)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
by: Lei, Ting, et al.
Published: (2024)
by: Lei, Ting, et al.
Published: (2024)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
by: Tang, Yingbo, et al.
Published: (2025)
by: Tang, Yingbo, et al.
Published: (2025)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
Vocabuild: An Accessible Augmented Tangible Interface for Gamified Vocabulary Learning of Constructing Meaning
by: Hu, Siying, et al.
Published: (2025)
by: Hu, Siying, et al.
Published: (2025)
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
by: Zheng, Yinan, et al.
Published: (2026)
by: Zheng, Yinan, et al.
Published: (2026)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection
by: Jiang, Donghong, et al.
Published: (2026)
by: Jiang, Donghong, et al.
Published: (2026)
H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows
by: Zhang, Harry, et al.
Published: (2025)
by: Zhang, Harry, et al.
Published: (2025)
OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies
by: Chen, Runnan, et al.
Published: (2024)
by: Chen, Runnan, et al.
Published: (2024)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Analytic DAG Constraints for Differentiable DAG Learning
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation
by: Jiang, Haochen, et al.
Published: (2024)
by: Jiang, Haochen, et al.
Published: (2024)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
by: Lin, Tzu-Jung, et al.
Published: (2025)
by: Lin, Tzu-Jung, et al.
Published: (2025)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
by: Yang, Zaiquan, et al.
Published: (2025)
by: Yang, Zaiquan, et al.
Published: (2025)
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation
by: Tang, Huadong, et al.
Published: (2024)
by: Tang, Huadong, et al.
Published: (2024)
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
by: Liu, Zhuoman, et al.
Published: (2024)
by: Liu, Zhuoman, et al.
Published: (2024)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
by: Feng, Jiarui, et al.
Published: (2026)
by: Feng, Jiarui, et al.
Published: (2026)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
by: Zhu, Guoliang, et al.
Published: (2026)
by: Zhu, Guoliang, et al.
Published: (2026)
Similar Items
-
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025) -
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2026) -
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
by: Tong, Edmond, et al.
Published: (2024) -
GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
by: Shao, Yawen, et al.
Published: (2024) -
SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
by: Wang, Hanqing, et al.
Published: (2025)