PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shang-Ching, Tran, Van Nhiem, Chen, Wenkai, Cheng, Wei-Lun, Huang, Yen-Lin, Liao, I-Bin, Li, Yung-Hui, Zhang, Jianwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
by: Ling, Suhan, et al.
Published: (2024)
by: Ling, Suhan, et al.
Published: (2024)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
ToolEENet: Tool Affordance 6D Pose Estimation
by: Wang, Yunlong, et al.
Published: (2024)
by: Wang, Yunlong, et al.
Published: (2024)
Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes
by: Li, Maximilian Xiling, et al.
Published: (2025)
by: Li, Maximilian Xiling, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
by: Li, Jinming, et al.
Published: (2024)
by: Li, Jinming, et al.
Published: (2024)
OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
by: Yu, Qiaojun, et al.
Published: (2024)
by: Yu, Qiaojun, et al.
Published: (2024)
GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning
by: Li, Mingleyang, et al.
Published: (2026)
by: Li, Mingleyang, et al.
Published: (2026)
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
DORA: Object Affordance-Guided Reinforcement Learning for Dexterous Robotic Manipulation
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
by: Kong, Weijie, et al.
Published: (2026)
by: Kong, Weijie, et al.
Published: (2026)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
SpotLight: Robotic Scene Understanding through Interaction and Affordance Detection
by: Engelbracht, Tim, et al.
Published: (2024)
by: Engelbracht, Tim, et al.
Published: (2024)
RAIL: Robot Affordance Imagination with Large Language Models
by: Zhang, Ceng, et al.
Published: (2024)
by: Zhang, Ceng, et al.
Published: (2024)
Vision-Language Model-based Physical Reasoning for Robot Liquid Perception
by: Lai, Wenqiang, et al.
Published: (2024)
by: Lai, Wenqiang, et al.
Published: (2024)
23 DoF Grasping Policies from a Raw Point Cloud
by: Matak, Martin, et al.
Published: (2024)
by: Matak, Martin, et al.
Published: (2024)
A Novel Task-Driven Diffusion-Based Policy with Affordance Learning for Generalizable Manipulation of Articulated Objects
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Scene-agnostic Hierarchical Bimanual Task Planning via Visual Affordance Reasoning
by: Lee, Kwang Bin, et al.
Published: (2025)
by: Lee, Kwang Bin, et al.
Published: (2025)
Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments
by: Li, Xingyi, et al.
Published: (2025)
by: Li, Xingyi, et al.
Published: (2025)
RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds
by: Zhao, Bin, et al.
Published: (2025)
by: Zhao, Bin, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
by: Shao, Xinyu, et al.
Published: (2025)
by: Shao, Xinyu, et al.
Published: (2025)
OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation
by: Narasimhan, Siddarth, et al.
Published: (2024)
by: Narasimhan, Siddarth, et al.
Published: (2024)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
by: Yan, Yu, et al.
Published: (2024)
by: Yan, Yu, et al.
Published: (2024)
Learning Affordances at Inference-Time for Vision-Language-Action Models
by: Shah, Ameesh, et al.
Published: (2025)
by: Shah, Ameesh, et al.
Published: (2025)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
by: Tong, Edmond, et al.
Published: (2024)
by: Tong, Edmond, et al.
Published: (2024)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
by: Guo, Wenkai, et al.
Published: (2025)
by: Guo, Wenkai, et al.
Published: (2025)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
Sailing Through Point Clouds: Safe Navigation Using Point Cloud Based Control Barrier Functions
by: Dai, Bolun, et al.
Published: (2024)
by: Dai, Bolun, et al.
Published: (2024)
3D Branch Point Cloud Completion for Robotic Pruning in Apple Orchards
by: Qiu, Tian, et al.
Published: (2024)
by: Qiu, Tian, et al.
Published: (2024)
HGACNet: Hierarchical Graph Attention Network for Cross-Modal Point Cloud Completion
by: Zeng, Yadan, et al.
Published: (2025)
by: Zeng, Yadan, et al.
Published: (2025)
Flying on Point Clouds with Reinforcement Learning
by: Xu, Guangtong, et al.
Published: (2025)
by: Xu, Guangtong, et al.
Published: (2025)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
Efficient Global Navigational Planning in 3D Structures based on Point Cloud Tomography
by: Yang, Bowen, et al.
Published: (2024)
by: Yang, Bowen, et al.
Published: (2024)
DynaHull: Density-centric Dynamic Point Filtering in Point Clouds
by: Habibiroudkenar, Pejman, et al.
Published: (2024)
by: Habibiroudkenar, Pejman, et al.
Published: (2024)
PlaneHEC: Efficient Hand-Eye Calibration for Multi-view Robotic Arm via Any Point Cloud Plane Detection
by: Wang, Ye, et al.
Published: (2025)
by: Wang, Ye, et al.
Published: (2025)
Incremental Learning of Full-Pose Via-Point Movement Primitives on Riemannian Manifolds
by: Daab, Tilman, et al.
Published: (2023)
by: Daab, Tilman, et al.
Published: (2023)
A Modular Pneumatic Soft Gripper Design for Aerial Grasping and Landing
by: Cheung, Hiu Ching, et al.
Published: (2023)
by: Cheung, Hiu Ching, et al.
Published: (2023)
Similar Items
-
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise
by: Ling, Suhan, et al.
Published: (2024) -
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024) -
ToolEENet: Tool Affordance 6D Pose Estimation
by: Wang, Yunlong, et al.
Published: (2024) -
Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes
by: Li, Maximilian Xiling, et al.
Published: (2025) -
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)