3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Guoqin, Jia, Qingxuan, Huang, Zeyuan, Chen, Gang, Ji, Ning, Yao, Zhipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perceiving, Reasoning, Adapting: A Dual-Layer Framework for VLM-Guided Precision Robotic Manipulation
by: Jia, Qingxuan, et al.
Published: (2025)
by: Jia, Qingxuan, et al.
Published: (2025)
VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing
by: Tang, Guoqin, et al.
Published: (2026)
by: Tang, Guoqin, et al.
Published: (2026)
Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
by: Chen, Hongyi, et al.
Published: (2024)
by: Chen, Hongyi, et al.
Published: (2024)
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
by: Guo, Pengyuan, et al.
Published: (2026)
by: Guo, Pengyuan, et al.
Published: (2026)
Grounded Vision-Language Interpreter for Integrated Task and Motion Planning
by: Siburian, Jeremy, et al.
Published: (2025)
by: Siburian, Jeremy, et al.
Published: (2025)
VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing
by: Zong, Junyi, et al.
Published: (2026)
by: Zong, Junyi, et al.
Published: (2026)
Vision-Language Interpreter for Robot Task Planning
by: Shirai, Keisuke, et al.
Published: (2023)
by: Shirai, Keisuke, et al.
Published: (2023)
LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
by: Zhang, Xiaohan, et al.
Published: (2025)
by: Zhang, Xiaohan, et al.
Published: (2025)
Vision-Language-Policy Model for Dynamic Robot Task Planning
by: Wang, Jin, et al.
Published: (2025)
by: Wang, Jin, et al.
Published: (2025)
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
by: Yang, Yuyuan, et al.
Published: (2026)
by: Yang, Yuyuan, et al.
Published: (2026)
LLM-based Robot Task Planning with Exceptional Handling for General Purpose Service Robots
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025)
by: Lou, Zhichen, et al.
Published: (2025)
Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning
by: Źróbek, Karolina, et al.
Published: (2026)
by: Źróbek, Karolina, et al.
Published: (2026)
RobotDesignGPT: Automated Robot Design Synthesis using Vision Language Models
by: Sontakke, Nitish, et al.
Published: (2026)
by: Sontakke, Nitish, et al.
Published: (2026)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
Automated Planning Domain Inference for Task and Motion Planning
by: Huang, Jinbang, et al.
Published: (2024)
by: Huang, Jinbang, et al.
Published: (2024)
Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning
by: Battistini, Cristiano, et al.
Published: (2026)
by: Battistini, Cristiano, et al.
Published: (2026)
Language-Grounded Hierarchical Planning and Execution with Multi-Robot 3D Scene Graphs
by: Strader, Jared, et al.
Published: (2025)
by: Strader, Jared, et al.
Published: (2025)
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
by: Vineet Bhat, et al.
Published: (2025)
by: Vineet Bhat, et al.
Published: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
by: Deigmoeller, Joerg, et al.
Published: (2026)
by: Deigmoeller, Joerg, et al.
Published: (2026)
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
by: Huang, Yuzhi, et al.
Published: (2026)
by: Huang, Yuzhi, et al.
Published: (2026)
From Perception to Symbolic Task Planning: Vision-Language Guided Human-Robot Collaborative Structured Assembly
by: Chen, Yanyi, et al.
Published: (2026)
by: Chen, Yanyi, et al.
Published: (2026)
Visual-Language-Guided Task Planning for Horticultural Robots
by: Cuaran, Jose, et al.
Published: (2026)
by: Cuaran, Jose, et al.
Published: (2026)
LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task Planning
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
Task-oriented Robotic Manipulation with Vision Language Models
by: Guran, Nurhan Bulus, et al.
Published: (2024)
by: Guran, Nurhan Bulus, et al.
Published: (2024)
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
by: He, Zeyuan, et al.
Published: (2026)
by: He, Zeyuan, et al.
Published: (2026)
Hierarchical LLM-Based Multi-Agent Framework with Prompt Optimization for Multi-Robot Task Planning
by: Kawabe, Tomoya, et al.
Published: (2026)
by: Kawabe, Tomoya, et al.
Published: (2026)
Event-Driven Proactive Assistive Manipulation with Grounded Vision-Language Planning
by: Liu, Fengkai, et al.
Published: (2026)
by: Liu, Fengkai, et al.
Published: (2026)
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
by: Chen, Annie S., et al.
Published: (2024)
by: Chen, Annie S., et al.
Published: (2024)
From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation
by: Feng, Zeyuan, et al.
Published: (2025)
by: Feng, Zeyuan, et al.
Published: (2025)
Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models
by: Xia, Xingyu, et al.
Published: (2026)
by: Xia, Xingyu, et al.
Published: (2026)
Leveraging Pre-trained Large Language Models with Refined Prompting for Online Task and Motion Planning
by: Guo, Huihui, et al.
Published: (2025)
by: Guo, Huihui, et al.
Published: (2025)
A Comparison of Prompt Engineering Techniques for Task Planning and Execution in Service Robotics
by: Bode, Jonas, et al.
Published: (2024)
by: Bode, Jonas, et al.
Published: (2024)
Surgical Task Automation Using Actor-Critic Frameworks and Self-Supervised Imitation Learning
by: Liu, Jingshuai, et al.
Published: (2024)
by: Liu, Jingshuai, et al.
Published: (2024)
FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction Following
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Robodimm: A Physics-Grounded Framework for Automated Actuator Sizing in Scalable Modular Robots
by: Torres, J. L., et al.
Published: (2026)
by: Torres, J. L., et al.
Published: (2026)
UniPlan: Vision-Language Task Planning for Mobile Manipulation with Unified PDDL Formulation
by: Ye, Haoming, et al.
Published: (2026)
by: Ye, Haoming, et al.
Published: (2026)
Similar Items
-
Perceiving, Reasoning, Adapting: A Dual-Layer Framework for VLM-Guided Precision Robotic Manipulation
by: Jia, Qingxuan, et al.
Published: (2025) -
VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing
by: Tang, Guoqin, et al.
Published: (2026) -
Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts
by: Chen, Hongyi, et al.
Published: (2024) -
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
by: Guo, Pengyuan, et al.
Published: (2026) -
Grounded Vision-Language Interpreter for Integrated Task and Motion Planning
by: Siburian, Jeremy, et al.
Published: (2025)