CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xiaoqi, Xu, Lingyun, Zhang, Mingxu, Liu, Jiaming, Shen, Yan, Ponomarenko, Iaroslav, Xu, Jiahui, Heng, Liang, Huang, Siyuan, Zhang, Shanghang, Dong, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
di: Liu, Jiaming, et al.
Pubblicazione: (2024)
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
di: Li, Chenxuan, et al.
Pubblicazione: (2024)
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
di: Heng, Liang, et al.
Pubblicazione: (2025)
di: Heng, Liang, et al.
Pubblicazione: (2025)
SpatialBot: Precise Spatial Understanding with Vision Language Models
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
di: Zhang, Mingxu, et al.
Pubblicazione: (2025)
Object-Centric Instruction Augmentation for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
di: Pan, Mingjie, et al.
Pubblicazione: (2023)
di: Pan, Mingjie, et al.
Pubblicazione: (2023)
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
di: Ji, Yuheng, et al.
Pubblicazione: (2025)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
di: Wang, Boyang, et al.
Pubblicazione: (2026)
di: Wang, Boyang, et al.
Pubblicazione: (2026)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
di: Zhu, Juan, et al.
Pubblicazione: (2026)
di: Zhu, Juan, et al.
Pubblicazione: (2026)
Crayon ingestion
di: Kisho Noda, et al.
Pubblicazione: (2025)
di: Kisho Noda, et al.
Pubblicazione: (2025)
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
di: Tan, Huajie, et al.
Pubblicazione: (2025)
di: Tan, Huajie, et al.
Pubblicazione: (2025)
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
di: Wu, Kun, et al.
Pubblicazione: (2024)
di: Wu, Kun, et al.
Pubblicazione: (2024)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
di: Gu, Chenyang, et al.
Pubblicazione: (2025)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
di: Liu, Zhuoyang, et al.
Pubblicazione: (2025)
di: Liu, Zhuoyang, et al.
Pubblicazione: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
di: Zheng, Ying, et al.
Pubblicazione: (2024)
di: Zheng, Ying, et al.
Pubblicazione: (2024)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
di: Ju, Yuanchen, et al.
Pubblicazione: (2024)
di: Ju, Yuanchen, et al.
Pubblicazione: (2024)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
di: Tan, Huajie, et al.
Pubblicazione: (2026)
di: Tan, Huajie, et al.
Pubblicazione: (2026)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2025)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
Language-Grounded Decoupled Action Representation for Robotic Manipulation
di: Weng, Wuding, et al.
Pubblicazione: (2026)
di: Weng, Wuding, et al.
Pubblicazione: (2026)
RoboPearls: Editable Video Simulation for Robot Manipulation
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
RoboCAS: A Benchmark for Robotic Manipulation in Complex Object Arrangement Scenarios
di: Zheng, Liming, et al.
Pubblicazione: (2024)
di: Zheng, Liming, et al.
Pubblicazione: (2024)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
di: Liu, Enguang, et al.
Pubblicazione: (2025)
di: Liu, Enguang, et al.
Pubblicazione: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
di: Li, Yajie, et al.
Pubblicazione: (2026)
di: Li, Yajie, et al.
Pubblicazione: (2026)
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
di: Zhang, Rongyu, et al.
Pubblicazione: (2025)
BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection
di: Liu, Jiaming, et al.
Pubblicazione: (2022)
di: Liu, Jiaming, et al.
Pubblicazione: (2022)
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
di: Xu, Qinwen, et al.
Pubblicazione: (2026)
di: Xu, Qinwen, et al.
Pubblicazione: (2026)
Exploring Sparse Visual Prompt for Domain Adaptive Dense Prediction
di: Yang, Senqiao, et al.
Pubblicazione: (2023)
di: Yang, Senqiao, et al.
Pubblicazione: (2023)
CoLLaVO: Crayon Large Language and Vision mOdel
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
di: Liu, Zhuoyang, et al.
Pubblicazione: (2026)
di: Liu, Zhuoyang, et al.
Pubblicazione: (2026)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
di: Xu, Siyuan, et al.
Pubblicazione: (2026)
di: Xu, Siyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
di: Liu, Jiaming, et al.
Pubblicazione: (2024) -
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
di: Kim, Taewhan, et al.
Pubblicazione: (2024) -
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
di: Li, Chenxuan, et al.
Pubblicazione: (2024) -
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
di: Huang, Siyuan, et al.
Pubblicazione: (2024) -
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
di: Heng, Liang, et al.
Pubblicazione: (2025)