Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, He, Kong, Quyu, Xu, Kechun, Xia, Xunlong, Deng, Bing, Ye, Jieping, Xiong, Rong, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025)
by: Lou, Zhichen, et al.
Published: (2025)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024)
by: Xu, Kechun, et al.
Published: (2024)
Toward Embodiment Equivariant Vision-Language-Action Policy
by: Chen, Anzhe, et al.
Published: (2025)
by: Chen, Anzhe, et al.
Published: (2025)
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023)
by: Xu, Kechun, et al.
Published: (2023)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
by: Xie, Liang, et al.
Published: (2022)
by: Xie, Liang, et al.
Published: (2022)
Towards Unified Interactive Visual Grounding in The Wild
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
by: Yang, Yifei, et al.
Published: (2026)
by: Yang, Yifei, et al.
Published: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Learning Affordances from Interactive Exploration using an Object-level Map
by: Wulkop, Paula, et al.
Published: (2025)
by: Wulkop, Paula, et al.
Published: (2025)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Revisit Mixture Models for Multi-Agent Simulation: Experimental Study within a Unified Framework
by: Lin, Longzhong, et al.
Published: (2025)
by: Lin, Longzhong, et al.
Published: (2025)
BEV-DWPVO: BEV-based Differentiable Weighted Procrustes for Low Scale-drift Monocular Visual Odometry on Ground
by: Wei, Yufei, et al.
Published: (2025)
by: Wei, Yufei, et al.
Published: (2025)
Affordance-Aware Interactive Decision-Making and Execution for Ambiguous Instructions
by: Xu, Hengxuan, et al.
Published: (2026)
by: Xu, Hengxuan, et al.
Published: (2026)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
DORA: Object Affordance-Guided Reinforcement Learning for Dexterous Robotic Manipulation
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
by: Tong, Edmond, et al.
Published: (2024)
by: Tong, Edmond, et al.
Published: (2024)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
by: Li, Jinming, et al.
Published: (2024)
by: Li, Jinming, et al.
Published: (2024)
Multi-Object Graph Affordance Network: Goal-Oriented Planning through Learned Compound Object Affordances
by: Girgin, Tuba, et al.
Published: (2023)
by: Girgin, Tuba, et al.
Published: (2023)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
by: Zhu, Guoliang, et al.
Published: (2026)
by: Zhu, Guoliang, et al.
Published: (2026)
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
by: Jia, Wanjun, et al.
Published: (2025)
by: Jia, Wanjun, et al.
Published: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
by: Honerkamp, Daniel, et al.
Published: (2024)
by: Honerkamp, Daniel, et al.
Published: (2024)
BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
by: Wei, Yufei, et al.
Published: (2025)
by: Wei, Yufei, et al.
Published: (2025)
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
by: Wu, Shijie, et al.
Published: (2024)
by: Wu, Shijie, et al.
Published: (2024)
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
by: Huang, Siyuan, et al.
Published: (2024)
by: Huang, Siyuan, et al.
Published: (2024)
DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
by: Li, Quanzhou, et al.
Published: (2025)
by: Li, Quanzhou, et al.
Published: (2025)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
by: Kong, Weijie, et al.
Published: (2026)
by: Kong, Weijie, et al.
Published: (2026)
AeroPlace-Flow: Language-Grounded Object Placement for Aerial Manipulators via Visual Foresight and Object Flow
by: Mishra, Sarthak, et al.
Published: (2026)
by: Mishra, Sarthak, et al.
Published: (2026)
GauTOAO: Gaussian-based Task-Oriented Affordance of Objects
by: Wang, Jiawen, et al.
Published: (2024)
by: Wang, Jiawen, et al.
Published: (2024)
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
by: Ling, Suhan, et al.
Published: (2024)
by: Ling, Suhan, et al.
Published: (2024)
SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
by: Fang, Kuan, et al.
Published: (2025)
by: Fang, Kuan, et al.
Published: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
Learning Language-Conditioned Deformable Object Manipulation with Graph Dynamics
by: Deng, Yuhong, et al.
Published: (2023)
by: Deng, Yuhong, et al.
Published: (2023)
Observation Time Difference: an Online Dynamic Objects Removal Method for Ground Vehicles
by: Wu, Rongguang, et al.
Published: (2024)
by: Wu, Rongguang, et al.
Published: (2024)
VADER: Visual Affordance Detection and Error Recovery for Multi Robot Human Collaboration
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
Similar Items
-
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025) -
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025) -
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024) -
Toward Embodiment Equivariant Vision-Language-Action Policy
by: Chen, Anzhe, et al.
Published: (2025) -
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023)