Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Haonan, Guo, Jingxiang, Wang, Bangjun, Zhang, Tianrui, Huang, Xuchuan, Zheng, Boren, Hou, Yiwen, Tie, Chenrui, Deng, Jiajun, Shao, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
ShapeForce: Low-Cost Soft Robotic Wrist for Contact-Rich Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025)
by: Zhu, Jinxuan, et al.
Published: (2025)
AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025)
by: Zhu, Jinxuan, et al.
Published: (2025)
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
by: Tie, Chenrui, et al.
Published: (2025)
by: Tie, Chenrui, et al.
Published: (2025)
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
Disentangled Object-Centric Image Representation for Robotic Manipulation
by: Emukpere, David, et al.
Published: (2025)
by: Emukpere, David, et al.
Published: (2025)
ManiFoundation Model for General-Purpose Robotic Manipulation of Contact Synthesis with Arbitrary Objects and Robots
by: Xu, Zhixuan, et al.
Published: (2024)
by: Xu, Zhixuan, et al.
Published: (2024)
$\mathcal{D(R,O)}$ Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping
by: Wei, Zhenyu, et al.
Published: (2024)
by: Wei, Zhenyu, et al.
Published: (2024)
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
IG-RFT: An Interaction-Guided RL Framework for VLA Models in Long-Horizon Robotic Manipulation
by: Su, Zhian, et al.
Published: (2026)
by: Su, Zhian, et al.
Published: (2026)
Object-Centric World Model for Language-Guided Manipulation
by: Jeong, Youngjoon, et al.
Published: (2025)
by: Jeong, Youngjoon, et al.
Published: (2025)
LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
by: Guan, Tianrui, et al.
Published: (2024)
by: Guan, Tianrui, et al.
Published: (2024)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
by: Yang, Yue, et al.
Published: (2026)
by: Yang, Yue, et al.
Published: (2026)
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
by: Hanyu, Taisei, et al.
Published: (2025)
by: Hanyu, Taisei, et al.
Published: (2025)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)
by: Lin, Dixuan, et al.
Published: (2025)
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
by: Heng, Liang, et al.
Published: (2025)
by: Heng, Liang, et al.
Published: (2025)
SR-Nav: Spatial Relationships Matter for Zero-shot Object Goal Navigation
by: Fang, Leyuan, et al.
Published: (2026)
by: Fang, Leyuan, et al.
Published: (2026)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
Object-Centric Kinodynamic Planning for Nonprehensile Robot Rearrangement Manipulation
by: Ren, Kejia, et al.
Published: (2024)
by: Ren, Kejia, et al.
Published: (2024)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
Object-Centric Representations Improve Policy Generalization in Robot Manipulation
by: Chapin, Alexandre, et al.
Published: (2025)
by: Chapin, Alexandre, et al.
Published: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating
by: Chen, Junting, et al.
Published: (2025)
by: Chen, Junting, et al.
Published: (2025)
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2024)
by: Raychaudhuri, Sonia, et al.
Published: (2024)
Learning Part-Aware Dense 3D Feature Field for Generalizable Articulated Object Manipulation
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
by: Fan, Cunxin, et al.
Published: (2025)
by: Fan, Cunxin, et al.
Published: (2025)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
by: Chung, Nhat, et al.
Published: (2025)
by: Chung, Nhat, et al.
Published: (2025)
UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
Empowering Small VLMs to Think with Dynamic Memorization and Exploration
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
by: Dai, Aobotao, et al.
Published: (2025)
by: Dai, Aobotao, et al.
Published: (2025)
Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation
by: Shen, Zhixuan, et al.
Published: (2026)
by: Shen, Zhixuan, et al.
Published: (2026)
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
by: Tian, Yuxuan, et al.
Published: (2026)
by: Tian, Yuxuan, et al.
Published: (2026)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation
by: Kedia, Kushal, et al.
Published: (2026)
by: Kedia, Kushal, et al.
Published: (2026)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
by: Xiao, Jianqiang, et al.
Published: (2025)
by: Xiao, Jianqiang, et al.
Published: (2025)
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
by: Yang, Tianshuo, et al.
Published: (2026)
by: Yang, Tianshuo, et al.
Published: (2026)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Similar Items
-
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
by: Chen, Zixuan, et al.
Published: (2026) -
ShapeForce: Low-Cost Soft Robotic Wrist for Contact-Rich Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025) -
AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
by: Zhu, Jinxuan, et al.
Published: (2025) -
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
by: Tie, Chenrui, et al.
Published: (2025) -
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
by: Unlu, Halil Utku, et al.
Published: (2024)