Pre-training Auto-regressive Robotic Models with 4D Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niu, Dantong, Sharma, Yuvan, Xue, Haoru, Biamby, Giscard, Zhang, Junyi, Ji, Ziteng, Darrell, Trevor, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
von: Yin, Yida, et al.
Veröffentlicht: (2024)
von: Yin, Yida, et al.
Veröffentlicht: (2024)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
von: Hsieh, Wen-Han, et al.
Veröffentlicht: (2025)
von: Hsieh, Wen-Han, et al.
Veröffentlicht: (2025)
Learning to Grasp Anything by Playing with Random Toys
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
From Generated Human Videos to Physically Plausible Robot Trajectories
von: Ni, James, et al.
Veröffentlicht: (2025)
von: Ni, James, et al.
Veröffentlicht: (2025)
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
von: Xue, Haoru, et al.
Veröffentlicht: (2025)
von: Xue, Haoru, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
AutoOdom: Learning Auto-regressive Proprioceptive Odometry for Legged Locomotion
von: Luo, Changsheng, et al.
Veröffentlicht: (2025)
von: Luo, Changsheng, et al.
Veröffentlicht: (2025)
VIP: Vision Instructed Pre-training for Robotic Manipulation
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
Spatiotemporal Predictive Pre-training for Robotic Motor Control
von: Yang, Jiange, et al.
Veröffentlicht: (2024)
von: Yang, Jiange, et al.
Veröffentlicht: (2024)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
von: Jiang, Guangqi, et al.
Veröffentlicht: (2024)
WROOM: An Autonomous Driving Approach for Off-Road Navigation
von: Kalaria, Dvij, et al.
Veröffentlicht: (2024)
von: Kalaria, Dvij, et al.
Veröffentlicht: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
von: Ding, Wei, et al.
Veröffentlicht: (2024)
von: Ding, Wei, et al.
Veröffentlicht: (2024)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Learning Model Predictive Control with Error Dynamics Regression for Autonomous Racing
von: Xue, Haoru, et al.
Veröffentlicht: (2023)
von: Xue, Haoru, et al.
Veröffentlicht: (2023)
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
von: Zheng, Ruijie, et al.
Veröffentlicht: (2026)
von: Zheng, Ruijie, et al.
Veröffentlicht: (2026)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing
von: Xue, Haoru, et al.
Veröffentlicht: (2024)
von: Xue, Haoru, et al.
Veröffentlicht: (2024)
AnyCar to Anywhere: Learning Universal Dynamics Model for Agile and Adaptive Mobility
von: Xiao, Wenli, et al.
Veröffentlicht: (2024)
von: Xiao, Wenli, et al.
Veröffentlicht: (2024)
Learning Humanoid Locomotion over Challenging Terrain
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
von: Radosavovic, Ilija, et al.
Veröffentlicht: (2024)
GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
von: Gao, Minquan, et al.
Veröffentlicht: (2025)
von: Gao, Minquan, et al.
Veröffentlicht: (2025)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
Pre-Trained Masked Image Model for Mobile Robot Navigation
von: Sharma, Vishnu Dutt, et al.
Veröffentlicht: (2023)
von: Sharma, Vishnu Dutt, et al.
Veröffentlicht: (2023)
Lifelike Agility and Play in Quadrupedal Robots using Reinforcement Learning and Generative Pre-trained Models
von: Han, Lei, et al.
Veröffentlicht: (2023)
von: Han, Lei, et al.
Veröffentlicht: (2023)
G$^{2}$TR: Generalized Grounded Temporal Reasoning for Robot Instruction Following by Combining Large Pre-trained Models
von: Arora, Riya, et al.
Veröffentlicht: (2024)
von: Arora, Riya, et al.
Veröffentlicht: (2024)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Where to Fetch: Extracting Visual Scene Representation from Large Pre-Trained Models for Robotic Goal Navigation
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
DML-RAM: Deep Multimodal Learning Framework for Robotic Arm Manipulation using Pre-trained Models
von: Kumar, Sathish, et al.
Veröffentlicht: (2025)
von: Kumar, Sathish, et al.
Veröffentlicht: (2025)
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
Unifying Deep Predicate Invention with Pre-trained Foundation Models
von: Wang, Qianwei, et al.
Veröffentlicht: (2025)
von: Wang, Qianwei, et al.
Veröffentlicht: (2025)
Pheno-Robot: An Auto-Digital Modelling System for In-Situ Phenotyping in the Field
von: Pan, Yaoqiang, et al.
Veröffentlicht: (2024)
von: Pan, Yaoqiang, et al.
Veröffentlicht: (2024)
Loopy Movements: Emergence of Rotation in a Multicellular Robot
von: Smith, Trevor, et al.
Veröffentlicht: (2024)
von: Smith, Trevor, et al.
Veröffentlicht: (2024)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024) -
In-Context Learning Enables Robot Action Prediction in LLMs
von: Yin, Yida, et al.
Veröffentlicht: (2024) -
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
von: Hsieh, Wen-Han, et al.
Veröffentlicht: (2025) -
Learning to Grasp Anything by Playing with Random Toys
von: Niu, Dantong, et al.
Veröffentlicht: (2025) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)