NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ran, Shen, Yan, Li, Xiaoqi, Wu, Ruihai, Dong, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
by: Wu, Ruihai, et al.
Published: (2025)
by: Wu, Ruihai, et al.
Published: (2025)
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
by: Wu, Ruihai, et al.
Published: (2023)
by: Wu, Ruihai, et al.
Published: (2023)
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
by: Chen, Yue, et al.
Published: (2024)
by: Chen, Yue, et al.
Published: (2024)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
by: Zhang, Xiaojie, et al.
Published: (2025)
by: Zhang, Xiaojie, et al.
Published: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
by: Shao, Rui, et al.
Published: (2025)
by: Shao, Rui, et al.
Published: (2025)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
by: Zhu, Xiaomeng, et al.
Published: (2025)
by: Zhu, Xiaomeng, et al.
Published: (2025)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
by: Jiang, Zebin, et al.
Published: (2025)
by: Jiang, Zebin, et al.
Published: (2025)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
by: Zhou, Dingyi, et al.
Published: (2026)
by: Zhou, Dingyi, et al.
Published: (2026)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
by: Tang, Yihe, et al.
Published: (2025)
by: Tang, Yihe, et al.
Published: (2025)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
by: Lan, Zihan, et al.
Published: (2025)
by: Lan, Zihan, et al.
Published: (2025)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
by: Waheed, Sania, et al.
Published: (2025)
by: Waheed, Sania, et al.
Published: (2025)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
by: Xiong, Chuyan, et al.
Published: (2024)
by: Xiong, Chuyan, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation
by: Lu, Haoran, et al.
Published: (2024)
by: Lu, Haoran, et al.
Published: (2024)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
by: Zhang, Mingxu, et al.
Published: (2025)
by: Zhang, Mingxu, et al.
Published: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
PhotoBot: Reference-Guided Interactive Photography via Natural Language
by: Limoyo, Oliver, et al.
Published: (2024)
by: Limoyo, Oliver, et al.
Published: (2024)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
by: Fan, Hongwei, et al.
Published: (2025)
by: Fan, Hongwei, et al.
Published: (2025)
ET-SEED: Efficient Trajectory-Level SE(3) Equivariant Diffusion Policy
by: Tie, Chenrui, et al.
Published: (2024)
by: Tie, Chenrui, et al.
Published: (2024)
Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding
by: Song, Yaoxian, et al.
Published: (2023)
by: Song, Yaoxian, et al.
Published: (2023)
Similar Items
-
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
by: Kim, Taewhan, et al.
Published: (2024) -
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
by: Wu, Ruihai, et al.
Published: (2025) -
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
by: Wu, Ruihai, et al.
Published: (2023) -
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
by: Chen, Yue, et al.
Published: (2024) -
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
by: Ding, Kairui, et al.
Published: (2024)