Saved in:
| Main Authors: | Li, Chengshu, Liang, Jacky, Zeng, Andy, Chen, Xinyun, Hausman, Karol, Sadigh, Dorsa, Levine, Sergey, Fei-Fei, Li, Xia, Fei, Ichter, Brian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2312.04474 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
GenCHiP: Generating Robot Policy Code for High-Precision and Contact-Rich Manipulation Tasks
by: Burns, Kaylee, et al.
Published: (2024)
by: Burns, Kaylee, et al.
Published: (2024)
Generative Expressive Robot Behaviors using Large Language Models
by: Mahadevan, Karthik, et al.
Published: (2024)
by: Mahadevan, Karthik, et al.
Published: (2024)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024)
by: Grannen, Jennifer, et al.
Published: (2024)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025)
by: Grannen, Jennifer, et al.
Published: (2025)
Grounding Robot Generalization in Training Data via Retrieval-Augmented VLMs
by: Gao, Jensen, et al.
Published: (2026)
by: Gao, Jensen, et al.
Published: (2026)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
by: Li, Haozhuo, et al.
Published: (2025)
by: Li, Haozhuo, et al.
Published: (2025)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies
by: Xie, Amber, et al.
Published: (2026)
by: Xie, Amber, et al.
Published: (2026)
CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments
by: Sathyamoorthy, Adarsh Jagan, et al.
Published: (2024)
by: Sathyamoorthy, Adarsh Jagan, et al.
Published: (2024)
Precise Robot Command Understanding Using Grammar-Constrained Large Language Models
by: Huo, Xinyun, et al.
Published: (2026)
by: Huo, Xinyun, et al.
Published: (2026)
RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
by: Gao, Tian, et al.
Published: (2026)
by: Gao, Tian, et al.
Published: (2026)
Joint Action Language Modelling for Transparent Policy Execution
by: Wulff, Theodor, et al.
Published: (2025)
by: Wulff, Theodor, et al.
Published: (2025)
Data Analogies Enable Efficient Cross-Embodiment Transfer
by: Yang, Jonathan, et al.
Published: (2026)
by: Yang, Jonathan, et al.
Published: (2026)
Toward Grounded Commonsense Reasoning
by: Kwon, Minae, et al.
Published: (2023)
by: Kwon, Minae, et al.
Published: (2023)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
by: Yang, Jonathan, et al.
Published: (2024)
by: Yang, Jonathan, et al.
Published: (2024)
Will People Enjoy a Robot Trainer? A Case Study with Snoopie the Pacerbot
by: Du, Maximilian, et al.
Published: (2026)
by: Du, Maximilian, et al.
Published: (2026)
Invariance Co-training for Robot Visual Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Language Guided Skill Discovery
by: Rho, Seungeun, et al.
Published: (2024)
by: Rho, Seungeun, et al.
Published: (2024)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Vision Language Models are In-Context Value Learners
by: Ma, Yecheng Jason, et al.
Published: (2024)
by: Ma, Yecheng Jason, et al.
Published: (2024)
Latent Diffusion Planning for Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
Data Retrieval with Importance Weights for Few-Shot Imitation Learning
by: Xie, Amber, et al.
Published: (2025)
by: Xie, Amber, et al.
Published: (2025)
What Matters for Batch Online Reinforcement Learning in Robotics?
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
by: Qian, Kangan, et al.
Published: (2025)
by: Qian, Kangan, et al.
Published: (2025)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
by: Tang, Yihe, et al.
Published: (2025)
by: Tang, Yihe, et al.
Published: (2025)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
IPPON: Common Sense Guided Informative Path Planning for Object Goal Navigation
by: Qu, Kaixian, et al.
Published: (2024)
by: Qu, Kaixian, et al.
Published: (2024)
Similar Items
-
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024) -
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023) -
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
by: Nasiriany, Soroush, et al.
Published: (2024) -
GenCHiP: Generating Robot Policy Code for High-Precision and Contact-Rich Manipulation Tasks
by: Burns, Kaylee, et al.
Published: (2024) -
Generative Expressive Robot Behaviors using Large Language Models
by: Mahadevan, Karthik, et al.
Published: (2024)