Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Dinesh, Sushil Samuel, Park, Shinkyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
by: Feng, Yunhai, et al.
Published: (2025)
by: Feng, Yunhai, et al.
Published: (2025)
λ: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics
by: Jaafar, Ahmed, et al.
Published: (2024)
by: Jaafar, Ahmed, et al.
Published: (2024)
RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
by: Jiang, Hanxiao, et al.
Published: (2024)
by: Jiang, Hanxiao, et al.
Published: (2024)
Autoregressive Action Sequence Learning for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models
by: Turcato, Niccolò, et al.
Published: (2025)
by: Turcato, Niccolò, et al.
Published: (2025)
Integrating Disambiguation and User Preferences into Large Language Models for Robot Motion Planning
by: Abugurain, Mohammed, et al.
Published: (2024)
by: Abugurain, Mohammed, et al.
Published: (2024)
Vision-Language Foundation Models as Effective Robot Imitators
by: Li, Xinghang, et al.
Published: (2023)
by: Li, Xinghang, et al.
Published: (2023)
VLMgineer: Vision Language Models as Robotic Toolsmiths
by: Gao, George Jiayuan, et al.
Published: (2025)
by: Gao, George Jiayuan, et al.
Published: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
by: Liu, Jinyi, et al.
Published: (2024)
by: Liu, Jinyi, et al.
Published: (2024)
Plan First, Diffuse Later: Extrinsic Graph Guidance for Long-Horizon Diffusion Planning
by: Hassidof, Yaniv, et al.
Published: (2026)
by: Hassidof, Yaniv, et al.
Published: (2026)
Learning Multi-Agent Loco-Manipulation for Long-Horizon Quadrupedal Pushing
by: Feng, Yuming, et al.
Published: (2024)
by: Feng, Yuming, et al.
Published: (2024)
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
by: Su, Haochen, et al.
Published: (2025)
by: Su, Haochen, et al.
Published: (2025)
Learning to Transfer Human Hand Skills for Robot Manipulations
by: Park, Sungjae, et al.
Published: (2025)
by: Park, Sungjae, et al.
Published: (2025)
Online Foundation Model Selection in Robotics
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Towards Interpretable Foundation Models of Robot Behavior: A Task Specific Policy Generation Approach
by: Sheidlower, Isaac, et al.
Published: (2024)
by: Sheidlower, Isaac, et al.
Published: (2024)
Learning Neuro-symbolic Programs for Language Guided Robot Manipulation
by: Kalithasan, Namasivayam, et al.
Published: (2022)
by: Kalithasan, Namasivayam, et al.
Published: (2022)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data
by: Tang, Grace, et al.
Published: (2024)
by: Tang, Grace, et al.
Published: (2024)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
by: Lu, Dekun, et al.
Published: (2025)
by: Lu, Dekun, et al.
Published: (2025)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
Eurekaverse: Environment Curriculum Generation via Large Language Models
by: Liang, William, et al.
Published: (2024)
by: Liang, William, et al.
Published: (2024)
Refining Compositional Diffusion for Reliable Long-Horizon Planning
by: Lee, Kyowoon, et al.
Published: (2026)
by: Lee, Kyowoon, et al.
Published: (2026)
Embodied Red Teaming for Auditing Robotic Foundation Models
by: Karnik, Sathwik, et al.
Published: (2024)
by: Karnik, Sathwik, et al.
Published: (2024)
Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
SLIM: Sim-to-Real Legged Instructive Manipulation via Long-Horizon Visuomotor Learning
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Geometric Red-Teaming for Robotic Manipulation
by: Goel, Divyam, et al.
Published: (2025)
by: Goel, Divyam, et al.
Published: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
Dynamics-Guided Diffusion Model for Sensor-less Robot Manipulator Design
by: Xu, Xiaomeng, et al.
Published: (2024)
by: Xu, Xiaomeng, et al.
Published: (2024)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
by: Xiong, Zheng, et al.
Published: (2025)
by: Xiong, Zheng, et al.
Published: (2025)
Localized Graph-Based Neural Dynamics Models for Terrain Manipulation
by: Liu, Chaoqi, et al.
Published: (2025)
by: Liu, Chaoqi, et al.
Published: (2025)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
by: Xue, Yuquan, et al.
Published: (2025)
by: Xue, Yuquan, et al.
Published: (2025)
Unsupervised Learning of Effective Actions in Robotics
by: Zaric, Marko, et al.
Published: (2024)
by: Zaric, Marko, et al.
Published: (2024)
Adaptive Diffusion Policy Optimization for Robotic Manipulation
by: Jiang, Huiyun, et al.
Published: (2025)
by: Jiang, Huiyun, et al.
Published: (2025)
Zero-Shot Visual Generalization in Robot Manipulation
by: Batra, Sumeet, et al.
Published: (2025)
by: Batra, Sumeet, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
Humanoid Manipulation Interface: Humanoid Whole-Body Manipulation from Robot-Free Demonstrations
by: Nai, Ruiqian, et al.
Published: (2026)
by: Nai, Ruiqian, et al.
Published: (2026)
Similar Items
-
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation
by: Feng, Yunhai, et al.
Published: (2025) -
λ: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics
by: Jaafar, Ahmed, et al.
Published: (2024) -
RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
by: Jiang, Hanxiao, et al.
Published: (2024) -
Autoregressive Action Sequence Learning for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024) -
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models
by: Turcato, Niccolò, et al.
Published: (2025)