Pre-training Auto-regressive Robotic Models with 4D Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Niu, Dantong, Sharma, Yuvan, Xue, Haoru, Biamby, Giscard, Zhang, Junyi, Ji, Ziteng, Darrell, Trevor, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024)
by: Yin, Yida, et al.
Published: (2024)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
Learning to Grasp Anything by Playing with Random Toys
by: Niu, Dantong, et al.
Published: (2025)
by: Niu, Dantong, et al.
Published: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
AutoOdom: Learning Auto-regressive Proprioceptive Odometry for Legged Locomotion
by: Luo, Changsheng, et al.
Published: (2025)
by: Luo, Changsheng, et al.
Published: (2025)
VIP: Vision Instructed Pre-training for Robotic Manipulation
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Spatiotemporal Predictive Pre-training for Robotic Motor Control
by: Yang, Jiange, et al.
Published: (2024)
by: Yang, Jiange, et al.
Published: (2024)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
WROOM: An Autonomous Driving Approach for Off-Road Navigation
by: Kalaria, Dvij, et al.
Published: (2024)
by: Kalaria, Dvij, et al.
Published: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
by: Ding, Wei, et al.
Published: (2024)
by: Ding, Wei, et al.
Published: (2024)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
by: Wu, Tsung-Han, et al.
Published: (2024)
by: Wu, Tsung-Han, et al.
Published: (2024)
Learning Model Predictive Control with Error Dynamics Regression for Autonomous Racing
by: Xue, Haoru, et al.
Published: (2023)
by: Xue, Haoru, et al.
Published: (2023)
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
by: Zheng, Ruijie, et al.
Published: (2026)
by: Zheng, Ruijie, et al.
Published: (2026)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
by: Li, Guangrun, et al.
Published: (2025)
by: Li, Guangrun, et al.
Published: (2025)
Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing
by: Xue, Haoru, et al.
Published: (2024)
by: Xue, Haoru, et al.
Published: (2024)
AnyCar to Anywhere: Learning Universal Dynamics Model for Agile and Adaptive Mobility
by: Xiao, Wenli, et al.
Published: (2024)
by: Xiao, Wenli, et al.
Published: (2024)
Learning Humanoid Locomotion over Challenging Terrain
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
by: Gao, Minquan, et al.
Published: (2025)
by: Gao, Minquan, et al.
Published: (2025)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025)
by: Li, Kelvin, et al.
Published: (2025)
Pre-Trained Masked Image Model for Mobile Robot Navigation
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
by: Sharma, Vishnu Dutt, et al.
Published: (2023)
Lifelike Agility and Play in Quadrupedal Robots using Reinforcement Learning and Generative Pre-trained Models
by: Han, Lei, et al.
Published: (2023)
by: Han, Lei, et al.
Published: (2023)
G$^{2}$TR: Generalized Grounded Temporal Reasoning for Robot Instruction Following by Combining Large Pre-trained Models
by: Arora, Riya, et al.
Published: (2024)
by: Arora, Riya, et al.
Published: (2024)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Where to Fetch: Extracting Visual Scene Representation from Large Pre-Trained Models for Robotic Goal Navigation
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
DML-RAM: Deep Multimodal Learning Framework for Robotic Arm Manipulation using Pre-trained Models
by: Kumar, Sathish, et al.
Published: (2025)
by: Kumar, Sathish, et al.
Published: (2025)
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Unifying Deep Predicate Invention with Pre-trained Foundation Models
by: Wang, Qianwei, et al.
Published: (2025)
by: Wang, Qianwei, et al.
Published: (2025)
Pheno-Robot: An Auto-Digital Modelling System for In-Situ Phenotyping in the Field
by: Pan, Yaoqiang, et al.
Published: (2024)
by: Pan, Yaoqiang, et al.
Published: (2024)
Loopy Movements: Emergence of Rotation in a Multicellular Robot
by: Smith, Trevor, et al.
Published: (2024)
by: Smith, Trevor, et al.
Published: (2024)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Similar Items
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024) -
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024) -
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025) -
Learning to Grasp Anything by Playing with Random Toys
by: Niu, Dantong, et al.
Published: (2025) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)