Video Generation with Learned Action Prior
Fuente:
arXiv
Saved in:
| Main Authors: | Sarkar, Meenakshi, Bhardwaj, Devansh, Ghose, Debasish |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Action-conditioned video data improves predictability
by: Sarkar, Meenakshi, et al.
Published: (2024)
by: Sarkar, Meenakshi, et al.
Published: (2024)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
by: Wang, Xinkai, et al.
Published: (2026)
by: Wang, Xinkai, et al.
Published: (2026)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)
by: Guo, Jun, et al.
Published: (2026)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors
by: Chen, Jiahe, et al.
Published: (2026)
by: Chen, Jiahe, et al.
Published: (2026)
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
by: Collins, Jeremy A., et al.
Published: (2025)
by: Collins, Jeremy A., et al.
Published: (2025)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
Learning Priors of Human Motion With Vision Transformers
by: Falqueto, Placido, et al.
Published: (2025)
by: Falqueto, Placido, et al.
Published: (2025)
G-DexGrasp: Generalizable Dexterous Grasping Synthesis Via Part-Aware Prior Retrieval and Prior-Assisted Generation
by: Jian, Juntao, et al.
Published: (2025)
by: Jian, Juntao, et al.
Published: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
by: Yan, Haodong, et al.
Published: (2026)
by: Yan, Haodong, et al.
Published: (2026)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Object Reconstruction under Occlusion with Generative Priors and Contact-induced Constraints
by: Zhu, Minghan, et al.
Published: (2025)
by: Zhu, Minghan, et al.
Published: (2025)
Dynamic Visual SLAM using a General 3D Prior
by: Zhong, Xingguang, et al.
Published: (2025)
by: Zhong, Xingguang, et al.
Published: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
by: Chen, Hanzhi, et al.
Published: (2025)
by: Chen, Hanzhi, et al.
Published: (2025)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
by: Liang, Junbang, et al.
Published: (2024)
by: Liang, Junbang, et al.
Published: (2024)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
by: Ma, Haoxiang, et al.
Published: (2024)
by: Ma, Haoxiang, et al.
Published: (2024)
Unifying Language-Action Understanding and Generation for Autonomous Driving
by: Wang, Xinyang, et al.
Published: (2026)
by: Wang, Xinyang, et al.
Published: (2026)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
by: Hui, Chenyu, et al.
Published: (2026)
by: Hui, Chenyu, et al.
Published: (2026)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
by: Wu, Xianjin, et al.
Published: (2026)
by: Wu, Xianjin, et al.
Published: (2026)
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026)
by: Su, Yue, et al.
Published: (2026)
SCAR: Self-Supervised Continuous Action Representation Learning
by: Liu, Hongjia, et al.
Published: (2026)
by: Liu, Hongjia, et al.
Published: (2026)
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
by: Lyu, Huaihai, et al.
Published: (2026)
by: Lyu, Huaihai, et al.
Published: (2026)
Exploring Real World Map Change Generalization of Prior-Informed HD Map Prediction Models
by: Bateman, Samuel M., et al.
Published: (2024)
by: Bateman, Samuel M., et al.
Published: (2024)
Learning Scene-Level Signed Directional Distance Function with Ellipsoidal Priors and Neural Residuals
by: Dai, Zhirui, et al.
Published: (2025)
by: Dai, Zhirui, et al.
Published: (2025)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Similar Items
-
Action-conditioned video data improves predictability
by: Sarkar, Meenakshi, et al.
Published: (2024) -
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025) -
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026) -
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026) -
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)