Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Kaizhe, Rui, Zihang, He, Yao, Liu, Yuyao, Hua, Pu, Xu, Huazhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo
by: Zhu, Junzhe, et al.
Published: (2024)
by: Zhu, Junzhe, et al.
Published: (2024)
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
by: Yuan, Zhecheng, et al.
Published: (2024)
by: Yuan, Zhecheng, et al.
Published: (2024)
AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
by: Zhang, Jiawei, et al.
Published: (2026)
by: Zhang, Jiawei, et al.
Published: (2026)
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
H$^3$DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization
by: Lei, Kun, et al.
Published: (2023)
by: Lei, Kun, et al.
Published: (2023)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
Diffusion Reward: Learning Rewards via Conditional Video Diffusion
by: Huang, Tao, et al.
Published: (2023)
by: Huang, Tao, et al.
Published: (2023)
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
ArrayBot: Reinforcement Learning for Generalizable Distributed Manipulation through Touch
by: Xue, Zhengrong, et al.
Published: (2023)
by: Xue, Zhengrong, et al.
Published: (2023)
On the Evaluation of Generative Robotic Simulations
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Towards Fusing Point Cloud and Visual Representations for Imitation Learning
by: Donat, Atalay, et al.
Published: (2025)
by: Donat, Atalay, et al.
Published: (2025)
EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
by: Punamiya, Ryan, et al.
Published: (2025)
by: Punamiya, Ryan, et al.
Published: (2025)
GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
by: Hua, Pu, et al.
Published: (2024)
by: Hua, Pu, et al.
Published: (2024)
Vision-based Xylem Wetness Classification in Stem Water Potential Determination
by: Peiris, Pamodya, et al.
Published: (2024)
by: Peiris, Pamodya, et al.
Published: (2024)
Adaptive Visual Imitation Learning for Robotic Assisted Feeding Across Varied Bowl Configurations and Food Types
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
by: Li, Huanyu, et al.
Published: (2026)
by: Li, Huanyu, et al.
Published: (2026)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
by: Yin, Zhenhan, et al.
Published: (2025)
by: Yin, Zhenhan, et al.
Published: (2025)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
by: Tian, Yufeng, et al.
Published: (2026)
by: Tian, Yufeng, et al.
Published: (2026)
VILP: Imitation Learning with Latent Video Planning
by: Xu, Zhengtong, et al.
Published: (2025)
by: Xu, Zhengtong, et al.
Published: (2025)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
Towards Learning a Generalizable 3D Scene Representation from 2D Observations
by: Gromniak, Martin, et al.
Published: (2026)
by: Gromniak, Martin, et al.
Published: (2026)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
by: Peng, Daojie, et al.
Published: (2026)
by: Peng, Daojie, et al.
Published: (2026)
Generalizable Image Repair for Robust Visual Control
by: Sobolewski, Carson, et al.
Published: (2025)
by: Sobolewski, Carson, et al.
Published: (2025)
DA-MMP: Learning Coordinated and Accurate Throwing with Dynamics-Aware Motion Manifold Primitives
by: Chu, Chi, et al.
Published: (2025)
by: Chu, Chi, et al.
Published: (2025)
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
by: Yuan, Zhecheng, et al.
Published: (2025)
by: Yuan, Zhecheng, et al.
Published: (2025)
Restoring Noisy Demonstration for Imitation Learning With Diffusion Models
by: Chen, Shang-Fu, et al.
Published: (2025)
by: Chen, Shang-Fu, et al.
Published: (2025)
Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
by: Hu, Kaizhe, et al.
Published: (2025)
by: Hu, Kaizhe, et al.
Published: (2025)
Generalizable Humanoid Manipulation with 3D Diffusion Policies
by: Ze, Yanjie, et al.
Published: (2024)
by: Ze, Yanjie, et al.
Published: (2024)
DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features
by: Yang, Jinghe, et al.
Published: (2025)
by: Yang, Jinghe, et al.
Published: (2025)
Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion
by: He, Honglin, et al.
Published: (2026)
by: He, Honglin, et al.
Published: (2026)
EgoMimic: Scaling Imitation Learning via Egocentric Video
by: Kareer, Simar, et al.
Published: (2024)
by: Kareer, Simar, et al.
Published: (2024)
FUNCTO: Function-Centric One-Shot Imitation Learning for Tool Manipulation
by: Tang, Chao, et al.
Published: (2025)
by: Tang, Chao, et al.
Published: (2025)
Learning Visual Quadrupedal Loco-Manipulation from Demonstrations
by: He, Zhengmao, et al.
Published: (2024)
by: He, Zhengmao, et al.
Published: (2024)
GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
by: Zhong, Yao, et al.
Published: (2025)
by: Zhong, Yao, et al.
Published: (2025)
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Similar Items
-
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024) -
DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo
by: Zhu, Junzhe, et al.
Published: (2024) -
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
by: Yuan, Zhecheng, et al.
Published: (2024) -
AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
by: Zhang, Jiawei, et al.
Published: (2026) -
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
by: Ze, Yanjie, et al.
Published: (2024)