Precise Action-to-Video Generation Through Visual Action Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuang, Wen, Chao, Guo, Haoyu, Peng, Sida, Qin, Minghan, Bao, Hujun, Zhou, Xiaowei, Hu, Ruizhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
by: Xia, Ziyuan, et al.
Published: (2026)
by: Xia, Ziyuan, et al.
Published: (2026)
SAM-guided Graph Cut for 3D Instance Segmentation
by: Guo, Haoyu, et al.
Published: (2023)
by: Guo, Haoyu, et al.
Published: (2023)
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
by: Cen, Zhi, et al.
Published: (2025)
by: Cen, Zhi, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Video Generation with Learned Action Prior
by: Sarkar, Meenakshi, et al.
Published: (2024)
by: Sarkar, Meenakshi, et al.
Published: (2024)
SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion
by: Guo, Xiyue, et al.
Published: (2025)
by: Guo, Xiyue, et al.
Published: (2025)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Open-Vocabulary Action Localization with Iterative Visual Prompting
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
World-Grounded Human Motion Recovery via Gravity-View Coordinates
by: Shen, Zehong, et al.
Published: (2024)
by: Shen, Zehong, et al.
Published: (2024)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
by: Li, Runhao, et al.
Published: (2025)
by: Li, Runhao, et al.
Published: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM
by: Zhang, Chuanrui, et al.
Published: (2026)
by: Zhang, Chuanrui, et al.
Published: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
by: Shi, Yiran, et al.
Published: (2026)
by: Shi, Yiran, et al.
Published: (2026)
Referring Atomic Video Action Recognition
by: Peng, Kunyu, et al.
Published: (2024)
by: Peng, Kunyu, et al.
Published: (2024)
Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
by: Li, Wenxin, et al.
Published: (2025)
by: Li, Wenxin, et al.
Published: (2025)
Latent Action Pretraining Through World Modeling
by: Tharwat, Bahey, et al.
Published: (2025)
by: Tharwat, Bahey, et al.
Published: (2025)
MaPa: Text-driven Photorealistic Material Painting for 3D Shapes
by: Zhang, Shangzan, et al.
Published: (2024)
by: Zhang, Shangzan, et al.
Published: (2024)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
by: Chen, Anthony, et al.
Published: (2025)
by: Chen, Anthony, et al.
Published: (2025)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
Latent Action Pretraining from Videos
by: Ye, Seonghyeon, et al.
Published: (2024)
by: Ye, Seonghyeon, et al.
Published: (2024)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
by: Li, Anqi, et al.
Published: (2025)
by: Li, Anqi, et al.
Published: (2025)
RD-VIO: Robust Visual-Inertial Odometry for Mobile Augmented Reality in Dynamic Environments
by: Li, Jinyu, et al.
Published: (2023)
by: Li, Jinyu, et al.
Published: (2023)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
by: Xiao, Wenli, et al.
Published: (2025)
by: Xiao, Wenli, et al.
Published: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
by: Zhou, Jiaming, et al.
Published: (2025)
by: Zhou, Jiaming, et al.
Published: (2025)
OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving
by: Wei, Julong, et al.
Published: (2024)
by: Wei, Julong, et al.
Published: (2024)
Similar Items
-
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
by: Xia, Ziyuan, et al.
Published: (2026) -
SAM-guided Graph Cut for 3D Instance Segmentation
by: Guo, Haoyu, et al.
Published: (2023) -
Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation
by: Cen, Zhi, et al.
Published: (2025) -
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026) -
Video Generation with Learned Action Prior
by: Sarkar, Meenakshi, et al.
Published: (2024)