ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Minjie, Zhu, Yichen, Li, Jinming, Zhou, Zhongyi, Wen, Junjie, Liu, Xiaoyu, Shen, Chaomin, Peng, Yaxin, Feng, Feifei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
Object-Centric Instruction Augmentation for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
di: Zhu, Minjie, et al.
Pubblicazione: (2024)
di: Zhu, Minjie, et al.
Pubblicazione: (2024)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
di: Zhu, Minjie, et al.
Pubblicazione: (2024)
di: Zhu, Minjie, et al.
Pubblicazione: (2024)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
di: Li, Jinming, et al.
Pubblicazione: (2024)
di: Li, Jinming, et al.
Pubblicazione: (2024)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
di: Li, Chengmeng, et al.
Pubblicazione: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
di: Wang, Wanying, et al.
Pubblicazione: (2024)
di: Wang, Wanying, et al.
Pubblicazione: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
di: Li, Yaxuan, et al.
Pubblicazione: (2025)
di: Li, Yaxuan, et al.
Pubblicazione: (2025)
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
di: Zhu, Yichen, et al.
Pubblicazione: (2025)
di: Zhu, Yichen, et al.
Pubblicazione: (2025)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
di: Zeng, Qiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Qiyuan, et al.
Pubblicazione: (2025)
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
di: Lim, Jia Syuen, et al.
Pubblicazione: (2026)
di: Lim, Jia Syuen, et al.
Pubblicazione: (2026)
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
di: Li, Jinming, et al.
Pubblicazione: (2024)
di: Li, Jinming, et al.
Pubblicazione: (2024)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
di: Wen, Junjie, et al.
Pubblicazione: (2025)
di: Wen, Junjie, et al.
Pubblicazione: (2025)
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
di: Wu, Kun, et al.
Pubblicazione: (2024)
di: Wu, Kun, et al.
Pubblicazione: (2024)
EgoPush: Learning End-to-End Egocentric Multi-Object Rearrangement for Mobile Robots
di: An, Boyuan, et al.
Pubblicazione: (2026)
di: An, Boyuan, et al.
Pubblicazione: (2026)
Dexterous Manipulation of Deformable Objects via Pneumatic Gripping: Lifting by One End
di: Mykhailyshyn, Roman, et al.
Pubblicazione: (2025)
di: Mykhailyshyn, Roman, et al.
Pubblicazione: (2025)
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
di: Kambara, Motonari, et al.
Pubblicazione: (2024)
di: Kambara, Motonari, et al.
Pubblicazione: (2024)
dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts
di: Jiang, Yuhua, et al.
Pubblicazione: (2026)
di: Jiang, Yuhua, et al.
Pubblicazione: (2026)
Vision-based Manipulation from Single Human Video with Open-World Object Graphs
di: Zhu, Yifeng, et al.
Pubblicazione: (2024)
di: Zhu, Yifeng, et al.
Pubblicazione: (2024)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
Latent Diffeomorphic Co-Design of End-Effectors for Deformable and Fragile Object Manipulation
di: Ikemura, Kei, et al.
Pubblicazione: (2026)
di: Ikemura, Kei, et al.
Pubblicazione: (2026)
Building Explicit World Model for Zero-Shot Open-World Object Manipulation
di: Li, Xiaotong, et al.
Pubblicazione: (2026)
di: Li, Xiaotong, et al.
Pubblicazione: (2026)
Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation
di: Devarakonda, Venkata Naren, et al.
Pubblicazione: (2024)
di: Devarakonda, Venkata Naren, et al.
Pubblicazione: (2024)
ScrewSplat: An End-to-End Method for Articulated Object Recognition
di: Kim, Seungyeon, et al.
Pubblicazione: (2025)
di: Kim, Seungyeon, et al.
Pubblicazione: (2025)
Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
di: Chen, Haonan, et al.
Pubblicazione: (2025)
di: Chen, Haonan, et al.
Pubblicazione: (2025)
YOPOv2-Tracker: An End-to-End Agile Tracking and Navigation Framework from Perception to Action
di: Lu, Junjie, et al.
Pubblicazione: (2025)
di: Lu, Junjie, et al.
Pubblicazione: (2025)
Efficient Feature Fusion for UAV Object Detection
di: Wang, Xudong, et al.
Pubblicazione: (2025)
di: Wang, Xudong, et al.
Pubblicazione: (2025)
Transferring Kinesthetic Demonstrations across Diverse Objects for Manipulation Planning
di: Das, Dibyendu, et al.
Pubblicazione: (2025)
di: Das, Dibyendu, et al.
Pubblicazione: (2025)
Generalizing End-To-End Autonomous Driving In Real-World Environments Using Zero-Shot LLMs
di: Dong, Zeyu, et al.
Pubblicazione: (2024)
di: Dong, Zeyu, et al.
Pubblicazione: (2024)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
di: Zhong, Rui, et al.
Pubblicazione: (2025)
di: Zhong, Rui, et al.
Pubblicazione: (2025)
Learning Generalizable Language-Conditioned Cloth Manipulation from Long Demonstrations
di: Zhao, Hanyi, et al.
Pubblicazione: (2025)
di: Zhao, Hanyi, et al.
Pubblicazione: (2025)
AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation
di: Wu, Yulu, et al.
Pubblicazione: (2025)
di: Wu, Yulu, et al.
Pubblicazione: (2025)
A Physics-informed End-to-End Occupancy Framework for Motion Planning of Autonomous Vehicles
di: Shen, Shuqi, et al.
Pubblicazione: (2025)
di: Shen, Shuqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
di: Wen, Junjie, et al.
Pubblicazione: (2024) -
Object-Centric Instruction Augmentation for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024) -
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025) -
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024) -
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
di: Wen, Junjie, et al.
Pubblicazione: (2025)