Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, He, Kong, Quyu, Xu, Kechun, Xia, Xunlong, Deng, Bing, Ye, Jieping, Xiong, Rong, Wang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
von: Lou, Zhichen, et al.
Veröffentlicht: (2025)
von: Lou, Zhichen, et al.
Veröffentlicht: (2025)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
von: Xu, Kechun, et al.
Veröffentlicht: (2024)
von: Xu, Kechun, et al.
Veröffentlicht: (2024)
Toward Embodiment Equivariant Vision-Language-Action Policy
von: Chen, Anzhe, et al.
Veröffentlicht: (2025)
von: Chen, Anzhe, et al.
Veröffentlicht: (2025)
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
von: Xu, Kechun, et al.
Veröffentlicht: (2023)
von: Xu, Kechun, et al.
Veröffentlicht: (2023)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
von: Xie, Liang, et al.
Veröffentlicht: (2022)
von: Xie, Liang, et al.
Veröffentlicht: (2022)
Towards Unified Interactive Visual Grounding in The Wild
von: Xu, Jie, et al.
Veröffentlicht: (2024)
von: Xu, Jie, et al.
Veröffentlicht: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
von: Yang, Yifei, et al.
Veröffentlicht: (2026)
von: Yang, Yifei, et al.
Veröffentlicht: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
Learning Affordances from Interactive Exploration using an Object-level Map
von: Wulkop, Paula, et al.
Veröffentlicht: (2025)
von: Wulkop, Paula, et al.
Veröffentlicht: (2025)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
von: Li, Jingliang, et al.
Veröffentlicht: (2026)
von: Li, Jingliang, et al.
Veröffentlicht: (2026)
Revisit Mixture Models for Multi-Agent Simulation: Experimental Study within a Unified Framework
von: Lin, Longzhong, et al.
Veröffentlicht: (2025)
von: Lin, Longzhong, et al.
Veröffentlicht: (2025)
BEV-DWPVO: BEV-based Differentiable Weighted Procrustes for Low Scale-drift Monocular Visual Odometry on Ground
von: Wei, Yufei, et al.
Veröffentlicht: (2025)
von: Wei, Yufei, et al.
Veröffentlicht: (2025)
Affordance-Aware Interactive Decision-Making and Execution for Ambiguous Instructions
von: Xu, Hengxuan, et al.
Veröffentlicht: (2026)
von: Xu, Hengxuan, et al.
Veröffentlicht: (2026)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
DORA: Object Affordance-Guided Reinforcement Learning for Dexterous Robotic Manipulation
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
von: Tong, Edmond, et al.
Veröffentlicht: (2024)
von: Tong, Edmond, et al.
Veröffentlicht: (2024)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
von: Chu, Hengshuo, et al.
Veröffentlicht: (2025)
von: Chu, Hengshuo, et al.
Veröffentlicht: (2025)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
Multi-Object Graph Affordance Network: Goal-Oriented Planning through Learned Compound Object Affordances
von: Girgin, Tuba, et al.
Veröffentlicht: (2023)
von: Girgin, Tuba, et al.
Veröffentlicht: (2023)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
von: Zhu, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhu, Guoliang, et al.
Veröffentlicht: (2026)
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
von: Jia, Wanjun, et al.
Veröffentlicht: (2025)
von: Jia, Wanjun, et al.
Veröffentlicht: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
von: Honerkamp, Daniel, et al.
Veröffentlicht: (2024)
von: Honerkamp, Daniel, et al.
Veröffentlicht: (2024)
BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
von: Wei, Yufei, et al.
Veröffentlicht: (2025)
von: Wei, Yufei, et al.
Veröffentlicht: (2025)
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
von: Li, Quanzhou, et al.
Veröffentlicht: (2025)
von: Li, Quanzhou, et al.
Veröffentlicht: (2025)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
AeroPlace-Flow: Language-Grounded Object Placement for Aerial Manipulators via Visual Foresight and Object Flow
von: Mishra, Sarthak, et al.
Veröffentlicht: (2026)
von: Mishra, Sarthak, et al.
Veröffentlicht: (2026)
GauTOAO: Gaussian-based Task-Oriented Affordance of Objects
von: Wang, Jiawen, et al.
Veröffentlicht: (2024)
von: Wang, Jiawen, et al.
Veröffentlicht: (2024)
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
von: Ling, Suhan, et al.
Veröffentlicht: (2024)
von: Ling, Suhan, et al.
Veröffentlicht: (2024)
SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
von: Fang, Kuan, et al.
Veröffentlicht: (2025)
von: Fang, Kuan, et al.
Veröffentlicht: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
von: Xu, Ran, et al.
Veröffentlicht: (2024)
von: Xu, Ran, et al.
Veröffentlicht: (2024)
Learning Language-Conditioned Deformable Object Manipulation with Graph Dynamics
von: Deng, Yuhong, et al.
Veröffentlicht: (2023)
von: Deng, Yuhong, et al.
Veröffentlicht: (2023)
Observation Time Difference: an Online Dynamic Objects Removal Method for Ground Vehicles
von: Wu, Rongguang, et al.
Veröffentlicht: (2024)
von: Wu, Rongguang, et al.
Veröffentlicht: (2024)
VADER: Visual Affordance Detection and Error Recovery for Multi Robot Human Collaboration
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
von: Xu, Kechun, et al.
Veröffentlicht: (2025) -
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
von: Lou, Zhichen, et al.
Veröffentlicht: (2025) -
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
von: Xu, Kechun, et al.
Veröffentlicht: (2024) -
Toward Embodiment Equivariant Vision-Language-Action Policy
von: Chen, Anzhe, et al.
Veröffentlicht: (2025) -
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
von: Xu, Kechun, et al.
Veröffentlicht: (2023)