SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Junjie, Bai, Chenjia, He, Haoran, Xia, Wenke, Wang, Zhigang, Zhao, Bin, Li, Xiu, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024)
by: He, Haoran, et al.
Published: (2024)
Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation
by: Feng, Ruoxuan, et al.
Published: (2024)
by: Feng, Ruoxuan, et al.
Published: (2024)
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
by: Fan, Chenyou, et al.
Published: (2025)
by: Fan, Chenyou, et al.
Published: (2025)
Learning Manipulation by Predicting Interaction
by: Zeng, Jia, et al.
Published: (2024)
by: Zeng, Jia, et al.
Published: (2024)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)
by: Yao, Yuanqi, et al.
Published: (2025)
AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
by: Yang, Dejie, et al.
Published: (2025)
by: Yang, Dejie, et al.
Published: (2025)
KOI: Accelerating Online Imitation Learning via Hybrid Key-state Guidance
by: Lu, Jingxian, et al.
Published: (2024)
by: Lu, Jingxian, et al.
Published: (2024)
VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
by: Wang, Fangyuan, et al.
Published: (2026)
by: Wang, Fangyuan, et al.
Published: (2026)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
by: Wu, Pengyuan, et al.
Published: (2026)
by: Wu, Pengyuan, et al.
Published: (2026)
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
by: Xu, Xiuwei, et al.
Published: (2024)
by: Xu, Xiuwei, et al.
Published: (2024)
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs
by: Xia, Wenke, et al.
Published: (2023)
by: Xia, Wenke, et al.
Published: (2023)
VLP: Vision-Language Preference Learning for Embodied Manipulation
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot Manipulation
by: Oh, Hanbit, et al.
Published: (2025)
by: Oh, Hanbit, et al.
Published: (2025)
Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
by: Wang, Jianzong, et al.
Published: (2026)
by: Wang, Jianzong, et al.
Published: (2026)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
by: Liu, Enguang, et al.
Published: (2025)
by: Liu, Enguang, et al.
Published: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
by: Xu, Wenjiang, et al.
Published: (2025)
by: Xu, Wenjiang, et al.
Published: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
MiMo-Embodied: X-Embodied Foundation Model Technical Report
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
by: Li, Puhao, et al.
Published: (2024)
by: Li, Puhao, et al.
Published: (2024)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
by: Zheng, Jiakun, et al.
Published: (2026)
by: Zheng, Jiakun, et al.
Published: (2026)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
by: Dong, Runpei, et al.
Published: (2026)
by: Dong, Runpei, et al.
Published: (2026)
Closed Loop Interactive Embodied Reasoning for Robot Manipulation
by: Nazarczuk, Michal, et al.
Published: (2024)
by: Nazarczuk, Michal, et al.
Published: (2024)
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
by: Yang, Tianshuo, et al.
Published: (2026)
by: Yang, Tianshuo, et al.
Published: (2026)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
Towards Fusing Point Cloud and Visual Representations for Imitation Learning
by: Donat, Atalay, et al.
Published: (2025)
by: Donat, Atalay, et al.
Published: (2025)
Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
by: Zhang, Xiaojie, et al.
Published: (2025)
by: Zhang, Xiaojie, et al.
Published: (2025)
FUNCTO: Function-Centric One-Shot Imitation Learning for Tool Manipulation
by: Tang, Chao, et al.
Published: (2025)
by: Tang, Chao, et al.
Published: (2025)
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
Universal Actions for Enhanced Embodied Foundation Models
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
by: Cui, Wenbo, et al.
Published: (2025)
by: Cui, Wenbo, et al.
Published: (2025)
Similar Items
-
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024) -
Play to the Score: Stage-Guided Dynamic Multi-Sensory Fusion for Robotic Manipulation
by: Feng, Ruoxuan, et al.
Published: (2024) -
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
by: Fan, Chenyou, et al.
Published: (2025) -
Learning Manipulation by Predicting Interaction
by: Zeng, Jia, et al.
Published: (2024) -
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)