ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Shaofei, Mu, Zhancun, Liu, Anji, Liang, Yitao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
von: Cai, Shaofei, et al.
Veröffentlicht: (2024)
von: Cai, Shaofei, et al.
Veröffentlicht: (2024)
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
von: Lian, Kewei, et al.
Veröffentlicht: (2025)
von: Lian, Kewei, et al.
Veröffentlicht: (2025)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
von: Tian, Ran, et al.
Veröffentlicht: (2024)
von: Tian, Ran, et al.
Veröffentlicht: (2024)
Open-World Skill Discovery from Unsegmented Demonstrations
von: Deng, Jingwen, et al.
Veröffentlicht: (2025)
von: Deng, Jingwen, et al.
Veröffentlicht: (2025)
FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
von: Zhong, Yiming, et al.
Veröffentlicht: (2025)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2025)
ImitDiff: Transferring Foundation-Model Priors for Distraction Robust Visuomotor Policy
von: Dong, Yuhang, et al.
Veröffentlicht: (2025)
von: Dong, Yuhang, et al.
Veröffentlicht: (2025)
View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
von: Tian, Stephen, et al.
Veröffentlicht: (2024)
von: Tian, Stephen, et al.
Veröffentlicht: (2024)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
SLIM: Sim-to-Real Legged Instructive Manipulation via Long-Horizon Visuomotor Learning
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
von: Zhang, Haichao, et al.
Veröffentlicht: (2025)
H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
von: Lu, Yiyang, et al.
Veröffentlicht: (2025)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2024)
von: Cai, Shaofei, et al.
Veröffentlicht: (2024)
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
von: Seo, Junwon, et al.
Veröffentlicht: (2026)
von: Seo, Junwon, et al.
Veröffentlicht: (2026)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
Smart Help: Strategic Opponent Modeling for Proactive and Adaptive Robot Assistance in Households
von: Cao, Zhihao, et al.
Veröffentlicht: (2024)
von: Cao, Zhihao, et al.
Veröffentlicht: (2024)
Video Diffusion Alignment via Reward Gradients
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence
von: Gupta, Pranay, et al.
Veröffentlicht: (2025)
von: Gupta, Pranay, et al.
Veröffentlicht: (2025)
Extrapolated Urban View Synthesis Benchmark
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
von: He, Haoyang, et al.
Veröffentlicht: (2025)
von: He, Haoyang, et al.
Veröffentlicht: (2025)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
von: Tran, Chi-Nguyen, et al.
Veröffentlicht: (2026)
von: Tran, Chi-Nguyen, et al.
Veröffentlicht: (2026)
Instant Policy: In-Context Imitation Learning via Graph Diffusion
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
von: Vosylius, Vitalis, et al.
Veröffentlicht: (2024)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
von: Blank, Nils, et al.
Veröffentlicht: (2024)
von: Blank, Nils, et al.
Veröffentlicht: (2024)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
von: Ze, Yanjie, et al.
Veröffentlicht: (2024)
von: Ze, Yanjie, et al.
Veröffentlicht: (2024)
Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
von: Gillman, Nate, et al.
Veröffentlicht: (2026)
von: Gillman, Nate, et al.
Veröffentlicht: (2026)
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2026)
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2026)
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
von: Liu, Songming, et al.
Veröffentlicht: (2026)
von: Liu, Songming, et al.
Veröffentlicht: (2026)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
von: Chen, Xiao, et al.
Veröffentlicht: (2024)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
VoxAct-B: Voxel-Based Acting and Stabilizing Policy for Bimanual Manipulation
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2024)
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2024)
FACTR: Force-Attending Curriculum Training for Contact-Rich Policy Learning
von: Liu, Jason Jingzhou, et al.
Veröffentlicht: (2025)
von: Liu, Jason Jingzhou, et al.
Veröffentlicht: (2025)
Planning with the Views via Scene Self-Exploration
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023)
von: Man, Yunze, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
von: Cai, Shaofei, et al.
Veröffentlicht: (2024) -
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025) -
LoopNav: Benchmarking Spatial Consistency in World Models
von: Lian, Kewei, et al.
Veröffentlicht: (2025) -
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
von: Tian, Ran, et al.
Veröffentlicht: (2024) -
Open-World Skill Discovery from Unsegmented Demonstrations
von: Deng, Jingwen, et al.
Veröffentlicht: (2025)