ViPRA: Video Prediction for Robot Actions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Routray, Sandeep, Pan, Hengkai, Jain, Unnat, Bahl, Shikhar, Pathak, Deepak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
von: Patel, Shivansh, et al.
Veröffentlicht: (2025)
von: Patel, Shivansh, et al.
Veröffentlicht: (2025)
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
von: Lin, Leo, et al.
Veröffentlicht: (2026)
von: Lin, Leo, et al.
Veröffentlicht: (2026)
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
von: Zhao, Zhida, et al.
Veröffentlicht: (2025)
von: Zhao, Zhida, et al.
Veröffentlicht: (2025)
HRP: Human Affordances for Robotic Pre-Training
von: Srirama, Mohan Kumar, et al.
Veröffentlicht: (2024)
von: Srirama, Mohan Kumar, et al.
Veröffentlicht: (2024)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
Video Diffusion Alignment via Reward Gradients
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
3D-VLA: A 3D Vision-Language-Action Generative World Model
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
von: Lian, Shijie, et al.
Veröffentlicht: (2026)
von: Lian, Shijie, et al.
Veröffentlicht: (2026)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills
von: Wang, Yunfei, et al.
Veröffentlicht: (2026)
von: Wang, Yunfei, et al.
Veröffentlicht: (2026)
Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
von: Kåsene, Vebjørn Haug, et al.
Veröffentlicht: (2025)
von: Kåsene, Vebjørn Haug, et al.
Veröffentlicht: (2025)
Learning from Massive Human Videos for Universal Humanoid Pose Control
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
von: Schroeder, Philip, et al.
Veröffentlicht: (2025)
von: Schroeder, Philip, et al.
Veröffentlicht: (2025)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
von: Yuan, Puzhen, et al.
Veröffentlicht: (2025)
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
von: Liu, Yibin, et al.
Veröffentlicht: (2026)
von: Liu, Yibin, et al.
Veröffentlicht: (2026)
Towards Predicting Any Human Trajectory In Context
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
von: Fujii, Ryo, et al.
Veröffentlicht: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
von: Bartoccioni, Florent, et al.
Veröffentlicht: (2025)
von: Bartoccioni, Florent, et al.
Veröffentlicht: (2025)
FLAME: Learning to Navigate with Multimodal LLM in Urban Environments
von: Xu, Yunzhe, et al.
Veröffentlicht: (2024)
von: Xu, Yunzhe, et al.
Veröffentlicht: (2024)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
von: Park, Sung-Yeon, et al.
Veröffentlicht: (2025)
Coaching a Robotic Sonographer: Learning Robotic Ultrasound with Sparse Expert's Feedback
von: Raina, Deepak, et al.
Veröffentlicht: (2024)
von: Raina, Deepak, et al.
Veröffentlicht: (2024)
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control
von: Cui, Zichen Jeff, et al.
Veröffentlicht: (2024)
von: Cui, Zichen Jeff, et al.
Veröffentlicht: (2024)
Generative Image as Action Models
von: Shridhar, Mohit, et al.
Veröffentlicht: (2024)
von: Shridhar, Mohit, et al.
Veröffentlicht: (2024)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
von: Feng, Yao, et al.
Veröffentlicht: (2025)
von: Feng, Yao, et al.
Veröffentlicht: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
von: Patel, Shivansh, et al.
Veröffentlicht: (2025) -
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
von: Lin, Leo, et al.
Veröffentlicht: (2026) -
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
von: Zhao, Zhida, et al.
Veröffentlicht: (2025) -
HRP: Human Affordances for Robotic Pre-Training
von: Srirama, Mohan Kumar, et al.
Veröffentlicht: (2024) -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)