DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yang, Yang, Liudi, Eskandar, George, Shen, Fengyi, Altillawi, Mohammad, Liu, Ziyuan, Kutyniok, Gitta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
by: Eskandar, George, et al.
Published: (2026)
by: Eskandar, George, et al.
Published: (2026)
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
by: Yang, Liudi, et al.
Published: (2026)
by: Yang, Liudi, et al.
Published: (2026)
CE-NPBG: Connectivity Enhanced Neural Point-Based Graphics for Novel View Synthesis in Autonomous Driving Scenes
by: Altillawi, Mohammad, et al.
Published: (2025)
by: Altillawi, Mohammad, et al.
Published: (2025)
Lifelong 3D Mapping Framework for Hand-held & Robot-mounted LiDAR Mapping Systems
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
ConsistentDreamer: View-Consistent Meshes Through Balanced Multi-View Gaussian Optimization
by: Şahin, Onat, et al.
Published: (2025)
by: Şahin, Onat, et al.
Published: (2025)
Lightweight Learning from Actuation-Space Demonstrations via Flow Matching for Whole-Body Soft Robotic Grasping
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
LiDAR Loop Closure Detection using Semantic Graphs with Graph Attention Networks
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
Physics-Informed Video Diffusion For Shallow Water Equations
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Joint Flow Trajectory Optimization For Feasible Robot Motion Generation from Video Demonstrations
by: Dong, Xiaoxiang, et al.
Published: (2025)
by: Dong, Xiaoxiang, et al.
Published: (2025)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
DITTO: Demonstration Imitation by Trajectory Transformation
by: Heppert, Nick, et al.
Published: (2024)
by: Heppert, Nick, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Self-supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System
by: Dengxiong, Xiwen, et al.
Published: (2024)
by: Dengxiong, Xiwen, et al.
Published: (2024)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
by: Niu, Yida, et al.
Published: (2026)
by: Niu, Yida, et al.
Published: (2026)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
by: Zhou, Huayi, et al.
Published: (2025)
by: Zhou, Huayi, et al.
Published: (2025)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
by: Bai, Kaixin, et al.
Published: (2024)
by: Bai, Kaixin, et al.
Published: (2024)
DeFM: Learning Foundation Representations from Depth for Robotics
by: Patel, Manthan, et al.
Published: (2026)
by: Patel, Manthan, et al.
Published: (2026)
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-assisted Endoscopic Submucosal Dissection
by: Xu, Mengya, et al.
Published: (2024)
by: Xu, Mengya, et al.
Published: (2024)
RecNet: An Invertible Point Cloud Encoding through Range Image Embeddings for Multi-Robot Map Sharing and Reconstruction
by: Stathoulopoulos, Nikolaos, et al.
Published: (2024)
by: Stathoulopoulos, Nikolaos, et al.
Published: (2024)
KineDepth: Utilizing Robot Kinematics for Online Metric Depth Estimation
by: Atar, Soofiyan, et al.
Published: (2024)
by: Atar, Soofiyan, et al.
Published: (2024)
EndoMUST: Monocular Depth Estimation for Robotic Endoscopy via End-to-end Multi-step Self-supervised Training
by: Shao, Liangjing, et al.
Published: (2025)
by: Shao, Liangjing, et al.
Published: (2025)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
by: Zeng, Qiyuan, et al.
Published: (2025)
by: Zeng, Qiyuan, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
by: Jiang, Yuming, et al.
Published: (2025)
by: Jiang, Yuming, et al.
Published: (2025)
Learning Generalizable 3D Manipulation With 10 Demonstrations
by: Ren, Yu, et al.
Published: (2024)
by: Ren, Yu, et al.
Published: (2024)
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
by: Salloom, Tony, et al.
Published: (2025)
by: Salloom, Tony, et al.
Published: (2025)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
Cosmos-H-Surgical: Learning Surgical Robot Policies from Videos via World Modeling
by: He, Yufan, et al.
Published: (2025)
by: He, Yufan, et al.
Published: (2025)
High-Quality, ROS Compatible Video Encoding and Decoding for High-Definition Datasets
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation
by: Zheng, Jiamin, et al.
Published: (2026)
by: Zheng, Jiamin, et al.
Published: (2026)
Depth Restoration of Hand-Held Transparent Objects for Human-to-Robot Handover
by: Yu, Ran, et al.
Published: (2024)
by: Yu, Ran, et al.
Published: (2024)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
MM-ACT: Learn from Multimodal Parallel Generation to Act
by: Liang, Haotian, et al.
Published: (2025)
by: Liang, Haotian, et al.
Published: (2025)
Similar Items
-
RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping
by: Bai, Yang, et al.
Published: (2025) -
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
by: Yang, Liudi, et al.
Published: (2025) -
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
by: Eskandar, George, et al.
Published: (2026) -
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
by: Yang, Liudi, et al.
Published: (2025) -
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
by: Yang, Liudi, et al.
Published: (2026)