The Ingredients for Robotic Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Dasari, Sudeep, Mees, Oier, Zhao, Sebastian, Srirama, Mohan Kumar, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024)
by: Chen, William, et al.
Published: (2024)
HRP: Human Affordances for Robotic Pre-Training
by: Srirama, Mohan Kumar, et al.
Published: (2024)
by: Srirama, Mohan Kumar, et al.
Published: (2024)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
by: Pai, Jonas, et al.
Published: (2025)
by: Pai, Jonas, et al.
Published: (2025)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
by: Blank, Nils, et al.
Published: (2024)
by: Blank, Nils, et al.
Published: (2024)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
by: Huang, Chenguang, et al.
Published: (2025)
by: Huang, Chenguang, et al.
Published: (2025)
GHIL-Glue: Hierarchical Control with Filtered Subgoal Images
by: Hatch, Kyle B., et al.
Published: (2024)
by: Hatch, Kyle B., et al.
Published: (2024)
DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
by: Tao, Tony, et al.
Published: (2025)
by: Tao, Tony, et al.
Published: (2025)
Bimanual Dexterity for Complex Tasks
by: Shaw, Kenneth, et al.
Published: (2024)
by: Shaw, Kenneth, et al.
Published: (2024)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024)
by: Doshi, Ria, et al.
Published: (2024)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
Learning Visuotactile Skills with Two Multifingered Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
CRAFT: Video Diffusion for Bimanual Robot Data Generation
by: Chen, Jason, et al.
Published: (2026)
by: Chen, Jason, et al.
Published: (2026)
Semantically Controllable Augmentations for Generalizable Robot Learning
by: Chen, Zoey, et al.
Published: (2024)
by: Chen, Zoey, et al.
Published: (2024)
Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation
by: Ma, Xiao, et al.
Published: (2024)
by: Ma, Xiao, et al.
Published: (2024)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
by: Mohan, Riya, et al.
Published: (2026)
by: Mohan, Riya, et al.
Published: (2026)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation
by: Lancaster, Patrick, et al.
Published: (2023)
by: Lancaster, Patrick, et al.
Published: (2023)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
by: Hurtado, Juana Valeria, et al.
Published: (2025)
by: Hurtado, Juana Valeria, et al.
Published: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
by: Liang, Yichao, et al.
Published: (2024)
by: Liang, Yichao, et al.
Published: (2024)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
by: Hu, Yafei, et al.
Published: (2023)
by: Hu, Yafei, et al.
Published: (2023)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
by: Liu, Peiqi, et al.
Published: (2024)
by: Liu, Peiqi, et al.
Published: (2024)
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
by: Jangir, Yash, et al.
Published: (2025)
by: Jangir, Yash, et al.
Published: (2025)
On the Evaluation of Generative Robotic Simulations
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Neural Fields in Robotics: A Survey
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
STT: Stateful Tracking with Transformers for Autonomous Driving
by: Jing, Longlong, et al.
Published: (2024)
by: Jing, Longlong, et al.
Published: (2024)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
by: Ke, Tsung-Wei, et al.
Published: (2024)
by: Ke, Tsung-Wei, et al.
Published: (2024)
Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware Learning
by: Mohan, Rohit, et al.
Published: (2025)
by: Mohan, Rohit, et al.
Published: (2025)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
Bifurcation Identification for Ultrasound-driven Robotic Cannulation
by: Morales, Cecilia G., et al.
Published: (2024)
by: Morales, Cecilia G., et al.
Published: (2024)
Redundancy-aware Action Spaces for Robot Learning
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
Rapid Motor Adaptation for Robotic Manipulator Arms
by: Liang, Yichao, et al.
Published: (2023)
by: Liang, Yichao, et al.
Published: (2023)
Unlocking Generalization for Robotics via Modularity and Scale
by: Dalal, Murtaza
Published: (2025)
by: Dalal, Murtaza
Published: (2025)
Similar Items
-
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024) -
HRP: Human Affordances for Robotic Pre-Training
by: Srirama, Mohan Kumar, et al.
Published: (2024) -
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
by: Pai, Jonas, et al.
Published: (2025) -
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
by: Blank, Nils, et al.
Published: (2024) -
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
by: Huang, Chenguang, et al.
Published: (2025)