RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jiaming, Liu, Mengzhen, Wang, Zhenyu, An, Pengju, Li, Xiaoqi, Zhou, Kaichen, Yang, Senqiao, Zhang, Renrui, Guo, Yandong, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
von: Liu, Jiaming, et al.
Veröffentlicht: (2023)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
von: Wei, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wei, Xiaobao, et al.
Veröffentlicht: (2023)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
von: Wu, Kun, et al.
Veröffentlicht: (2024)
von: Wu, Kun, et al.
Veröffentlicht: (2024)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuoyang, et al.
Veröffentlicht: (2026)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
von: Pan, Mingjie, et al.
Veröffentlicht: (2023)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
von: Gu, Chenyang, et al.
Veröffentlicht: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation
von: Ni, Jiayi, et al.
Veröffentlicht: (2023)
von: Ni, Jiayi, et al.
Veröffentlicht: (2023)
Exploring Sparse Visual Prompt for Domain Adaptive Dense Prediction
von: Yang, Senqiao, et al.
Veröffentlicht: (2023)
von: Yang, Senqiao, et al.
Veröffentlicht: (2023)
BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection
von: Liu, Jiaming, et al.
Veröffentlicht: (2022)
von: Liu, Jiaming, et al.
Veröffentlicht: (2022)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies
von: Yuan, Chengbo, et al.
Veröffentlicht: (2025)
von: Yuan, Chengbo, et al.
Veröffentlicht: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
von: Wang, Guokang, et al.
Veröffentlicht: (2024)
von: Wang, Guokang, et al.
Veröffentlicht: (2024)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
von: Wu, Shihan, et al.
Veröffentlicht: (2025)
von: Wu, Shihan, et al.
Veröffentlicht: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2026)
Survey of Vision-Language-Action Models for Embodied Manipulation
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
von: Shao, Rui, et al.
Veröffentlicht: (2025)
von: Shao, Rui, et al.
Veröffentlicht: (2025)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024) -
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025) -
Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023) -
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
von: Liu, Jiaming, et al.
Veröffentlicht: (2023) -
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)