From Generated Human Videos to Physically Plausible Robot Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, James, Wang, Zekai, Lin, Wei, Bar, Amir, LeCun, Yann, Darrell, Trevor, Malik, Jitendra, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024)
by: Yin, Yida, et al.
Published: (2024)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Parallel Stochastic Gradient-Based Planning for World Models
by: Psenka, Michael, et al.
Published: (2026)
by: Psenka, Michael, et al.
Published: (2026)
OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
by: Dupoux, Emmanuel, et al.
Published: (2026)
by: Dupoux, Emmanuel, et al.
Published: (2026)
Pre-training Auto-regressive Robotic Models with 4D Representations
by: Niu, Dantong, et al.
Published: (2025)
by: Niu, Dantong, et al.
Published: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
Learning Humanoid Locomotion over Challenging Terrain
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
by: Zhou, Gaoyue, et al.
Published: (2024)
by: Zhou, Gaoyue, et al.
Published: (2024)
What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
by: Terver, Basile, et al.
Published: (2025)
by: Terver, Basile, et al.
Published: (2025)
Value-guided action planning with JEPA world models
by: Destrade, Matthieu, et al.
Published: (2025)
by: Destrade, Matthieu, et al.
Published: (2025)
Learning to Grasp Anything by Playing with Random Toys
by: Niu, Dantong, et al.
Published: (2025)
by: Niu, Dantong, et al.
Published: (2025)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Closing the Train-Test Gap in World Models for Gradient-Based Planning
by: Parthasarathy, Arjun, et al.
Published: (2025)
by: Parthasarathy, Arjun, et al.
Published: (2025)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024)
by: Hansen, Nicklas, et al.
Published: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment
by: Wi, Youngsun, et al.
Published: (2026)
by: Wi, Youngsun, et al.
Published: (2026)
Fast and Exact Enumeration of Deep Networks Partitions Regions
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
MonoDuo: Using One Robot Arm to Learn Bimanual Policies
by: Bajamahal, Sandeep, et al.
Published: (2026)
by: Bajamahal, Sandeep, et al.
Published: (2026)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation
by: Pan, Jiahe, et al.
Published: (2026)
by: Pan, Jiahe, et al.
Published: (2026)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Generalizing Robot Trajectories from Single-Context Human Demonstrations: A Probabilistic Approach
by: Lee, Qian Ying, et al.
Published: (2025)
by: Lee, Qian Ying, et al.
Published: (2025)
OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer
by: Yin, Jessica, et al.
Published: (2025)
by: Yin, Jessica, et al.
Published: (2025)
Enhancing Cognitive Robotics with Commonsense through LLM-Generated Preconditions and Subgoals
by: Bachner, Ohad, et al.
Published: (2025)
by: Bachner, Ohad, et al.
Published: (2025)
Similar Items
-
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025) -
Navigation World Models
by: Bar, Amir, et al.
Published: (2024) -
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024) -
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
by: Bar, Amir, et al.
Published: (2024) -
Parallel Stochastic Gradient-Based Planning for World Models
by: Psenka, Michael, et al.
Published: (2026)