GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
Fuente:
arXiv
Saved in:
| Main Authors: | Wake, Naoki, Kanehira, Atsushi, Sasabuchi, Kazuhiro, Takamatsu, Jun, Ikeuchi, Katsushi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-driven Behavior Tree for Context-aware Task Planning
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
Open-Vocabulary Action Localization with Iterative Visual Prompting
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
A Taxonomy of Self-Handover
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
IK Seed Generator for Dual-Arm Human-like Physicality Robot with Mobile Base
by: Takamatsu, Jun, et al.
Published: (2025)
by: Takamatsu, Jun, et al.
Published: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
by: Kanehira, Atsushi, et al.
Published: (2025)
by: Kanehira, Atsushi, et al.
Published: (2025)
Designing Library of Skill-Agents for Hardware-Level Reusability
by: Takamatsu, Jun, et al.
Published: (2024)
by: Takamatsu, Jun, et al.
Published: (2024)
APriCoT: Action Primitives based on Contact-state Transition for In-Hand Tool Manipulation
by: Saito, Daichi, et al.
Published: (2024)
by: Saito, Daichi, et al.
Published: (2024)
Harnessing GPT-4V(ision) for Insurance: A Preliminary Exploration
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023)
by: Chen, Yi, et al.
Published: (2023)
Signs of Language: Embodied Sign Language Fingerspelling Acquisition from Demonstrations for Human-Robot Interaction
by: Tavella, Federico, et al.
Published: (2022)
by: Tavella, Federico, et al.
Published: (2022)
Virtual Community: An Open World for Humans, Robots, and Society
by: Zhou, Qinhong, et al.
Published: (2025)
by: Zhou, Qinhong, et al.
Published: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
GPT-4V(ision) Unsuitable for Clinical Care and Education: A Clinician-Evaluated Assessment
by: Senkaiahliyan, Senthujan, et al.
Published: (2023)
by: Senkaiahliyan, Senthujan, et al.
Published: (2023)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
by: Zeng, Qiyuan, et al.
Published: (2025)
by: Zeng, Qiyuan, et al.
Published: (2025)
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
by: Hori, Chiori, et al.
Published: (2025)
by: Hori, Chiori, et al.
Published: (2025)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
by: Wang, Lirui, et al.
Published: (2023)
by: Wang, Lirui, et al.
Published: (2023)
LLM-Grounded Dynamic Task Planning with Hierarchical Temporal Logic for Human-Aware Multi-Robot Collaboration
by: Hu, Shuyuan, et al.
Published: (2026)
by: Hu, Shuyuan, et al.
Published: (2026)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
by: Guo, Dingkun, et al.
Published: (2024)
by: Guo, Dingkun, et al.
Published: (2024)
Learning Compositional Behaviors from Demonstration and Language
by: Liu, Weiyu, et al.
Published: (2025)
by: Liu, Weiyu, et al.
Published: (2025)
RefAV: Towards Planning-Centric Scenario Mining
by: Davidson, Cainan, et al.
Published: (2025)
by: Davidson, Cainan, et al.
Published: (2025)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
by: Wan, Zifu, et al.
Published: (2025)
by: Wan, Zifu, et al.
Published: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
by: Zhang, Shiduo, et al.
Published: (2024)
by: Zhang, Shiduo, et al.
Published: (2024)
Multimodal Anomaly Detection for Human-Robot Interaction
by: Ribeiro, Guilherme, et al.
Published: (2026)
by: Ribeiro, Guilherme, et al.
Published: (2026)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
by: An, Dong, et al.
Published: (2023)
by: An, Dong, et al.
Published: (2023)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)
by: Lu, Jinghui, et al.
Published: (2026)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026)
by: Schroeder, Philip, et al.
Published: (2026)
Similar Items
-
VLM-driven Behavior Tree for Context-aware Task Planning
by: Wake, Naoki, et al.
Published: (2025) -
Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
by: Sasabuchi, Kazuhiro, et al.
Published: (2025) -
Open-Vocabulary Action Localization with Iterative Visual Prompting
by: Wake, Naoki, et al.
Published: (2024) -
A Taxonomy of Self-Handover
by: Wake, Naoki, et al.
Published: (2025) -
Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience
by: Wake, Naoki, et al.
Published: (2024)