ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Siyuan, Ponomarenko, Iaroslav, Jiang, Zhengkai, Li, Xiaoqi, Hu, Xiaobin, Gao, Peng, Li, Hongsheng, Dong, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
von: Kim, Taewhan, et al.
Veröffentlicht: (2024)
von: Kim, Taewhan, et al.
Veröffentlicht: (2024)
BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration
von: Shen, Yan, et al.
Veröffentlicht: (2026)
von: Shen, Yan, et al.
Veröffentlicht: (2026)
UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024)
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints
von: Pan, Mingjie, et al.
Veröffentlicht: (2025)
von: Pan, Mingjie, et al.
Veröffentlicht: (2025)
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
von: Tang, Weiliang, et al.
Veröffentlicht: (2025)
von: Tang, Weiliang, et al.
Veröffentlicht: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
von: Xu, Ran, et al.
Veröffentlicht: (2024)
von: Xu, Ran, et al.
Veröffentlicht: (2024)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
von: Li, Kailin, et al.
Veröffentlicht: (2025)
von: Li, Kailin, et al.
Veröffentlicht: (2025)
iManip: Skill-Incremental Learning for Robotic Manipulation
von: Zheng, Zexin, et al.
Veröffentlicht: (2025)
von: Zheng, Zexin, et al.
Veröffentlicht: (2025)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
von: Tong, Edmond, et al.
Veröffentlicht: (2024)
von: Tong, Edmond, et al.
Veröffentlicht: (2024)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
UniDoorManip: Learning Universal Door Manipulation Policy Over Large-scale and Diverse Door Manipulation Environments
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
von: Shao, Xinyu, et al.
Veröffentlicht: (2025)
von: Shao, Xinyu, et al.
Veröffentlicht: (2025)
HeteroGenManip: Generalizable Manipulation For Heterogeneous Object Interactions
von: Shen, Zhenhao, et al.
Veröffentlicht: (2026)
von: Shen, Zhenhao, et al.
Veröffentlicht: (2026)
AffordanceLLM: Grounding Affordance from Vision Language Models
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
von: Kang, Zeyi, et al.
Veröffentlicht: (2025)
von: Kang, Zeyi, et al.
Veröffentlicht: (2025)
Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning
von: Wang, Yuanfei, et al.
Veröffentlicht: (2025)
von: Wang, Yuanfei, et al.
Veröffentlicht: (2025)
A3VLM: Actionable Articulation-Aware Vision Language Model
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
Learning Instruction-Guided Manipulation Affordance via Large Models for Embodied Robotic Tasks
von: Li, Dayou, et al.
Veröffentlicht: (2024)
von: Li, Dayou, et al.
Veröffentlicht: (2024)
RAIL: Robot Affordance Imagination with Large Language Models
von: Zhang, Ceng, et al.
Veröffentlicht: (2024)
von: Zhang, Ceng, et al.
Veröffentlicht: (2024)
Manip4Care: Robotic Manipulation of Human Limbs for Solving Assistive Tasks
von: Koh, Yubin, et al.
Veröffentlicht: (2025)
von: Koh, Yubin, et al.
Veröffentlicht: (2025)
BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly
von: Shen, Yan, et al.
Veröffentlicht: (2025)
von: Shen, Yan, et al.
Veröffentlicht: (2025)
FSAG: Enhancing Human-to-Dexterous-Hand Finger-Specific Affordance Grounding via Diffusion Models
von: Han, Yifan, et al.
Veröffentlicht: (2026)
von: Han, Yifan, et al.
Veröffentlicht: (2026)
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
von: Chen, Zijia, et al.
Veröffentlicht: (2026)
von: Chen, Zijia, et al.
Veröffentlicht: (2026)
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
von: Huang, Chengyue, et al.
Veröffentlicht: (2026)
von: Huang, Chengyue, et al.
Veröffentlicht: (2026)
UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph
von: Liu, Haichao, et al.
Veröffentlicht: (2026)
von: Liu, Haichao, et al.
Veröffentlicht: (2026)
RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation
von: An, Boshi, et al.
Veröffentlicht: (2023)
von: An, Boshi, et al.
Veröffentlicht: (2023)
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
von: Dong, Zibin, et al.
Veröffentlicht: (2025)
von: Dong, Zibin, et al.
Veröffentlicht: (2025)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
von: Zhao, Enyu, et al.
Veröffentlicht: (2025)
von: Zhao, Enyu, et al.
Veröffentlicht: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
von: Li, Puhao, et al.
Veröffentlicht: (2024)
von: Li, Puhao, et al.
Veröffentlicht: (2024)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation
von: Zhao, Ziyan, et al.
Veröffentlicht: (2025)
von: Zhao, Ziyan, et al.
Veröffentlicht: (2025)
Articulated-Body Dynamics Network: Dynamics-Grounded Prior for Robot Learning
von: Shin, Sangwoo, et al.
Veröffentlicht: (2026)
von: Shin, Sangwoo, et al.
Veröffentlicht: (2026)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
von: Sun, Yu, et al.
Veröffentlicht: (2026)
von: Sun, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
von: Kim, Taewhan, et al.
Veröffentlicht: (2024) -
BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration
von: Shen, Yan, et al.
Veröffentlicht: (2026) -
UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
von: Yu, Qiaojun, et al.
Veröffentlicht: (2024) -
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025) -
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
von: Li, Xin, et al.
Veröffentlicht: (2024)