mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pai, Jonas, Achenbach, Liam, Montesinos, Victoriano, Forrai, Benedek, Mees, Oier, Nava, Elvis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity
von: Nava, Elvis, et al.
Veröffentlicht: (2025)
von: Nava, Elvis, et al.
Veröffentlicht: (2025)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
von: Huang, Chenguang, et al.
Veröffentlicht: (2025)
von: Huang, Chenguang, et al.
Veröffentlicht: (2025)
Large Video Planner Enables Generalizable Robot Control
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
von: Blank, Nils, et al.
Veröffentlicht: (2024)
von: Blank, Nils, et al.
Veröffentlicht: (2024)
FASTER: Rethinking Real-Time Flow VLAs
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
World Model for Robot Learning: A Comprehensive Survey
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
von: Chen, Hanzhi, et al.
Veröffentlicht: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
von: Niu, Jiahui, et al.
Veröffentlicht: (2026)
von: Niu, Jiahui, et al.
Veröffentlicht: (2026)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
von: Fang, Yu, et al.
Veröffentlicht: (2026)
von: Fang, Yu, et al.
Veröffentlicht: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts
von: Yang, Yanting, et al.
Veröffentlicht: (2024)
von: Yang, Yanting, et al.
Veröffentlicht: (2024)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
von: Morin, Sacha, et al.
Veröffentlicht: (2025)
von: Morin, Sacha, et al.
Veröffentlicht: (2025)
Planar Velocity Estimation for Fast-Moving Mobile Robots Using Event-Based Optical Flow
von: Boyle, Liam, et al.
Veröffentlicht: (2025)
von: Boyle, Liam, et al.
Veröffentlicht: (2025)
Unified Video Action Model
von: Li, Shuang, et al.
Veröffentlicht: (2025)
von: Li, Shuang, et al.
Veröffentlicht: (2025)
World Knowledge from AI Image Generation for Robot Control
von: Krumme, Jonas, et al.
Veröffentlicht: (2025)
von: Krumme, Jonas, et al.
Veröffentlicht: (2025)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
von: Bharadhwaj, Homanga, et al.
Veröffentlicht: (2024)
von: Bharadhwaj, Homanga, et al.
Veröffentlicht: (2024)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
von: Zhao, Hongxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Hongxiang, et al.
Veröffentlicht: (2025)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
von: Wang, Siting, et al.
Veröffentlicht: (2026)
von: Wang, Siting, et al.
Veröffentlicht: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
Towards Generalizable Robotic Manipulation in Dynamic Environments
von: Fang, Heng, et al.
Veröffentlicht: (2026)
von: Fang, Heng, et al.
Veröffentlicht: (2026)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
von: Feng, Yixu, et al.
Veröffentlicht: (2026)
von: Feng, Yixu, et al.
Veröffentlicht: (2026)
Evaluating Real-World Robot Manipulation Policies in Simulation
von: Li, Xuanlin, et al.
Veröffentlicht: (2024)
von: Li, Xuanlin, et al.
Veröffentlicht: (2024)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
von: Nakamoto, Mitsuhiko, et al.
Veröffentlicht: (2024)
von: Nakamoto, Mitsuhiko, et al.
Veröffentlicht: (2024)
GHIL-Glue: Hierarchical Control with Filtered Subgoal Images
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
Semantically Controllable Augmentations for Generalizable Robot Learning
von: Chen, Zoey, et al.
Veröffentlicht: (2024)
von: Chen, Zoey, et al.
Veröffentlicht: (2024)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Sampling-Based Model Predictive Control for Dexterous Manipulation on a Biomimetic Tendon-Driven Hand
von: Hess, Adrian, et al.
Veröffentlicht: (2024)
von: Hess, Adrian, et al.
Veröffentlicht: (2024)
FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
von: Xu, Hongli, et al.
Veröffentlicht: (2025)
von: Xu, Hongli, et al.
Veröffentlicht: (2025)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
Generalizable Image Repair for Robust Visual Control
von: Sobolewski, Carson, et al.
Veröffentlicht: (2025)
von: Sobolewski, Carson, et al.
Veröffentlicht: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
DriveVA: Video Action Models are Zero-Shot Drivers
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
Precise Action-to-Video Generation Through Visual Action Prompts
von: Wang, Yuang, et al.
Veröffentlicht: (2025)
von: Wang, Yuang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity
von: Nava, Elvis, et al.
Veröffentlicht: (2025) -
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026) -
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024) -
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
von: Huang, Chenguang, et al.
Veröffentlicht: (2025) -
Large Video Planner Enables Generalizable Robot Control
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)