SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Haoquan, Grotz, Markus, Pumacay, Wilbert, Wang, Yi Ru, Fox, Dieter, Krishna, Ranjay, Duan, Jiafei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025)
von: Lee, Jason, et al.
Veröffentlicht: (2025)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
EVE: Enabling Anyone to Train Robots using Augmented Reality
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
von: Deshpande, Abhay, et al.
Veröffentlicht: (2026)
von: Deshpande, Abhay, et al.
Veröffentlicht: (2026)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
OptiGrasp: Optimized Grasp Pose Detection Using RGB Images for Warehouse Picking Robots
von: Atar, Soofiyan, et al.
Veröffentlicht: (2024)
von: Atar, Soofiyan, et al.
Veröffentlicht: (2024)
MolmoAct2: Action Reasoning Models for Real-world Deployment
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
von: Chen, Shirui, et al.
Veröffentlicht: (2026)
von: Chen, Shirui, et al.
Veröffentlicht: (2026)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
von: Kim, Yejin, et al.
Veröffentlicht: (2026)
von: Kim, Yejin, et al.
Veröffentlicht: (2026)
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
von: Yan, Ge, et al.
Veröffentlicht: (2025)
von: Yan, Ge, et al.
Veröffentlicht: (2025)
Active Vision for Scene Understanding
von: Grotz, Markus
Veröffentlicht: (2022)
von: Grotz, Markus
Veröffentlicht: (2022)
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
von: Sagar, Som, et al.
Veröffentlicht: (2024)
von: Sagar, Som, et al.
Veröffentlicht: (2024)
TetraGrip: Sensor-Driven Multi-Suction Reactive Object Manipulation in Cluttered Scenes
von: Torrado, Paolo, et al.
Veröffentlicht: (2025)
von: Torrado, Paolo, et al.
Veröffentlicht: (2025)
MemoAct: Atkinson-Shiffrin-Inspired Memory-Augmented Visuomotor Policy for Robotic Manipulation
von: Tan, Liufan, et al.
Veröffentlicht: (2026)
von: Tan, Liufan, et al.
Veröffentlicht: (2026)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
von: Zhang, Junjie, et al.
Veröffentlicht: (2024)
von: Zhang, Junjie, et al.
Veröffentlicht: (2024)
From Grounding to Manipulation: Case Studies of Foundation Model Integration in Embodied Robotic Systems
von: Sui, Xiuchao, et al.
Veröffentlicht: (2025)
von: Sui, Xiuchao, et al.
Veröffentlicht: (2025)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
von: Deshpande, Abhay, et al.
Veröffentlicht: (2025)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
von: Gu, Songen, et al.
Veröffentlicht: (2026)
von: Gu, Songen, et al.
Veröffentlicht: (2026)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
von: Muttaqien, Muhammad A., et al.
Veröffentlicht: (2025)
von: Muttaqien, Muhammad A., et al.
Veröffentlicht: (2025)
What Foundation Models can Bring for Robot Learning in Manipulation : A Survey
von: Li, Dingzhe, et al.
Veröffentlicht: (2024)
von: Li, Dingzhe, et al.
Veröffentlicht: (2024)
Transferring Foundation Models for Generalizable Robotic Manipulation
von: Yang, Jiange, et al.
Veröffentlicht: (2023)
von: Yang, Jiange, et al.
Veröffentlicht: (2023)
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
von: Lin, Chunru, et al.
Veröffentlicht: (2025)
von: Lin, Chunru, et al.
Veröffentlicht: (2025)
ASID: Active Exploration for System Identification in Robotic Manipulation
von: Memmel, Marius, et al.
Veröffentlicht: (2024)
von: Memmel, Marius, et al.
Veröffentlicht: (2024)
Innovative Integration of Visual Foundation Model with a Robotic Arm on a Mobile Platform
von: Zhang, Shimian, et al.
Veröffentlicht: (2024)
von: Zhang, Shimian, et al.
Veröffentlicht: (2024)
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?
von: Zhang, Zhongru, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongru, et al.
Veröffentlicht: (2026)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
von: Wang, Guokang, et al.
Veröffentlicht: (2024)
von: Wang, Guokang, et al.
Veröffentlicht: (2024)
RoboCade: Gamifying Robot Data Collection
von: Mirchandani, Suvir, et al.
Veröffentlicht: (2025)
von: Mirchandani, Suvir, et al.
Veröffentlicht: (2025)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
von: Bai, Shuanghao, et al.
Veröffentlicht: (2025)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2025)
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains
von: Wang, Yi Ru, et al.
Veröffentlicht: (2026)
von: Wang, Yi Ru, et al.
Veröffentlicht: (2026)
Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models
von: Tay, Tracey Yee Hsin, et al.
Veröffentlicht: (2026)
von: Tay, Tracey Yee Hsin, et al.
Veröffentlicht: (2026)
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024) -
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
von: Pumacay, Wilbert, et al.
Veröffentlicht: (2024) -
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
von: Duan, Jiafei, et al.
Veröffentlicht: (2024) -
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
von: Wang, Yi Ru, et al.
Veröffentlicht: (2025) -
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)