Prediction with Action: Visual Policy Learning via Joint Denoising Process
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yanjiang, Hu, Yucheng, Zhang, Jianke, Wang, Yen-Jen, Chen, Xiaoyu, Lu, Chaochao, Chen, Jianyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
by: Zhang, Jianke, et al.
Published: (2024)
by: Zhang, Jianke, et al.
Published: (2024)
Improving Vision-Language-Action Model with Online Reinforcement Learning
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023)
by: Guo, Yanjiang, et al.
Published: (2023)
Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
by: Gu, Xinyang, et al.
Published: (2024)
by: Gu, Xinyang, et al.
Published: (2024)
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
by: Zhang, Jianke, et al.
Published: (2025)
by: Zhang, Jianke, et al.
Published: (2025)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer
by: Gu, Xinyang, et al.
Published: (2024)
by: Gu, Xinyang, et al.
Published: (2024)
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?
by: Zhang, Zhongru, et al.
Published: (2026)
by: Zhang, Zhongru, et al.
Published: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
UAM: A Dual-Stream Perspective on Forgetting in VLA Training
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
Efficient and Generalized end-to-end Autonomous Driving System with Latent Deep Reinforcement Learning and Demonstrations
by: Tang, Zuojin, et al.
Published: (2024)
by: Tang, Zuojin, et al.
Published: (2024)
Learning Native Continuation for Action Chunking Flow Policies
by: Liu, Yufeng, et al.
Published: (2026)
by: Liu, Yufeng, et al.
Published: (2026)
ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
by: Gao, Jingyue, et al.
Published: (2026)
by: Gao, Jingyue, et al.
Published: (2026)
Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
Action with Visual Primitives
by: Guo, Weilong, et al.
Published: (2026)
by: Guo, Weilong, et al.
Published: (2026)
Human-assisted Robotic Policy Refinement via Action Preference Optimization
by: Xia, Wenke, et al.
Published: (2025)
by: Xia, Wenke, et al.
Published: (2025)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
by: Liufu, Weijia, et al.
Published: (2026)
by: Liufu, Weijia, et al.
Published: (2026)
Action-to-Action Flow Matching
by: Jia, Jindou, et al.
Published: (2026)
by: Jia, Jindou, et al.
Published: (2026)
SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization
by: Li, Jiashun, et al.
Published: (2026)
by: Li, Jiashun, et al.
Published: (2026)
A Contact-Safe Reinforcement Learning Framework for Contact-Rich Robot Manipulation
by: Zhu, Xiang, et al.
Published: (2022)
by: Zhu, Xiang, et al.
Published: (2022)
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
by: Xue, Han, et al.
Published: (2025)
by: Xue, Han, et al.
Published: (2025)
Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
by: Terence, Ng Wen Zheng, et al.
Published: (2024)
by: Terence, Ng Wen Zheng, et al.
Published: (2024)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
by: Hu, Xintong, et al.
Published: (2026)
by: Hu, Xintong, et al.
Published: (2026)
Traj2Action: A Co-Denoising Framework for Trajectory-Guided Human-to-Robot Skill Transfer
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Exploring Temporal Representation in Neural Processes for Multimodal Action Prediction
by: Fedozzi, Marco Gabriele, et al.
Published: (2026)
by: Fedozzi, Marco Gabriele, et al.
Published: (2026)
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
by: Chen, Xinzhe, et al.
Published: (2026)
by: Chen, Xinzhe, et al.
Published: (2026)
Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition
by: Quartey, Benedict, et al.
Published: (2026)
by: Quartey, Benedict, et al.
Published: (2026)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization
by: Chen, Zengjue, et al.
Published: (2025)
by: Chen, Zengjue, et al.
Published: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
by: Guo, Yanjiang, et al.
Published: (2026)
by: Guo, Yanjiang, et al.
Published: (2026)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
by: Ji, Kailun, et al.
Published: (2025)
by: Ji, Kailun, et al.
Published: (2025)
D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models
by: Guo, Yucheng, et al.
Published: (2026)
by: Guo, Yucheng, et al.
Published: (2026)
Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry
by: Yang, Zemin, et al.
Published: (2026)
by: Yang, Zemin, et al.
Published: (2026)
Coordinated Humanoid Manipulation with Choice Policies
by: Qi, Haozhi, et al.
Published: (2025)
by: Qi, Haozhi, et al.
Published: (2025)
Decision-Focused Learning to Predict Action Costs for Planning
by: Mandi, Jayanta, et al.
Published: (2024)
by: Mandi, Jayanta, et al.
Published: (2024)
Similar Items
-
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
by: Zhang, Jianke, et al.
Published: (2025) -
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024) -
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
by: Zhang, Jianke, et al.
Published: (2024) -
Improving Vision-Language-Action Model with Online Reinforcement Learning
by: Guo, Yanjiang, et al.
Published: (2025) -
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
by: Guo, Yanjiang, et al.
Published: (2023)