Video Generators are Robot Policies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Junbang, Tokmakov, Pavel, Liu, Ruoshi, Sudhakar, Sruthi, Shah, Paarth, Ambrus, Rares, Vondrick, Carl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
PaperBot: Learning to Design Real-World Tools Using Paper
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
Differentiable Robot Rendering
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
Understanding Video Transformers via Universal Concept Discovery
von: Kowal, Matthew, et al.
Veröffentlicht: (2024)
von: Kowal, Matthew, et al.
Veröffentlicht: (2024)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
Self-Improving Autonomous Underwater Manipulation
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
Controlling the World by Sleight of Hand
von: Sudhakar, Sruthi, et al.
Veröffentlicht: (2024)
von: Sudhakar, Sruthi, et al.
Veröffentlicht: (2024)
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
von: Guizilini, Vitor, et al.
Veröffentlicht: (2024)
von: Guizilini, Vitor, et al.
Veröffentlicht: (2024)
Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies
von: Xu, Chen, et al.
Veröffentlicht: (2025)
von: Xu, Chen, et al.
Veröffentlicht: (2025)
pix2gestalt: Amodal Segmentation by Synthesizing Wholes
von: Ozguroglu, Ege, et al.
Veröffentlicht: (2024)
von: Ozguroglu, Ege, et al.
Veröffentlicht: (2024)
Robot Learning as an Empirical Science: Best Practices for Policy Evaluation
von: Kress-Gazit, Hadas, et al.
Veröffentlicht: (2024)
von: Kress-Gazit, Hadas, et al.
Veröffentlicht: (2024)
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
von: Zhu, Chuning, et al.
Veröffentlicht: (2025)
von: Zhu, Chuning, et al.
Veröffentlicht: (2025)
Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
von: Yu, Justin, et al.
Veröffentlicht: (2025)
von: Yu, Justin, et al.
Veröffentlicht: (2025)
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
Capturing Visual Environment Structure Correlates with Control Performance
von: Dong, Jiahua, et al.
Veröffentlicht: (2026)
von: Dong, Jiahua, et al.
Veröffentlicht: (2026)
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
von: Snyder, David, et al.
Veröffentlicht: (2025)
von: Snyder, David, et al.
Veröffentlicht: (2025)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping
von: Iwase, Shun, et al.
Veröffentlicht: (2025)
von: Iwase, Shun, et al.
Veröffentlicht: (2025)
Neural Fields in Robotics: A Survey
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape
von: Wu, Rundi, et al.
Veröffentlicht: (2023)
von: Wu, Rundi, et al.
Veröffentlicht: (2023)
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
von: Vincent, Joseph A., et al.
Veröffentlicht: (2024)
von: Vincent, Joseph A., et al.
Veröffentlicht: (2024)
DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation
von: Ikeda, Takuya, et al.
Veröffentlicht: (2024)
von: Ikeda, Takuya, et al.
Veröffentlicht: (2024)
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
von: Zhu, Xiang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiang, et al.
Veröffentlicht: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
von: Yin, Guang, et al.
Veröffentlicht: (2025)
von: Yin, Guang, et al.
Veröffentlicht: (2025)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
von: Ye, Junjie, et al.
Veröffentlicht: (2026)
von: Ye, Junjie, et al.
Veröffentlicht: (2026)
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
von: Heng, Liang, et al.
Veröffentlicht: (2025)
von: Heng, Liang, et al.
Veröffentlicht: (2025)
CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects
von: Heppert, Nick, et al.
Veröffentlicht: (2023)
von: Heppert, Nick, et al.
Veröffentlicht: (2023)
A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation
von: Lin, Fanqi, et al.
Veröffentlicht: (2026)
von: Lin, Fanqi, et al.
Veröffentlicht: (2026)
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
von: Chaudhury, Arkadeep Narayan, et al.
Veröffentlicht: (2024)
von: Chaudhury, Arkadeep Narayan, et al.
Veröffentlicht: (2024)
New York Smells: A Large Multimodal Dataset for Olfaction
von: Ozguroglu, Ege, et al.
Veröffentlicht: (2025)
von: Ozguroglu, Ege, et al.
Veröffentlicht: (2025)
Policy Learning for Social Robot-Led Physiotherapy
von: Bettosi, Carl, et al.
Veröffentlicht: (2025)
von: Bettosi, Carl, et al.
Veröffentlicht: (2025)
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
von: Patil, Sarvesh, et al.
Veröffentlicht: (2026)
von: Patil, Sarvesh, et al.
Veröffentlicht: (2026)
Proximity and Visuotactile Point Cloud Fusion for Contact Patches in Extreme Deformation
von: Yin, Jessica, et al.
Veröffentlicht: (2023)
von: Yin, Jessica, et al.
Veröffentlicht: (2023)
DARE: Diffusion Policy for Autonomous Robot Exploration
von: Cao, Yuhong, et al.
Veröffentlicht: (2024)
von: Cao, Yuhong, et al.
Veröffentlicht: (2024)
EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images
von: Canberk, Alper, et al.
Veröffentlicht: (2024)
von: Canberk, Alper, et al.
Veröffentlicht: (2024)
MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
von: Shah, Rutav, et al.
Veröffentlicht: (2025)
von: Shah, Rutav, et al.
Veröffentlicht: (2025)
Masked Generative Policy for Robotic Control
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2025)
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2025)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
von: Liang, Junbang, et al.
Veröffentlicht: (2024) -
PaperBot: Learning to Design Real-World Tools Using Paper
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024) -
Differentiable Robot Rendering
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024) -
Understanding Video Transformers via Universal Concept Discovery
von: Kowal, Matthew, et al.
Veröffentlicht: (2024) -
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)