PhysiAgent: An Embodied Agent Framework in Physical World

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhihao, Li, Jianxiong, Zheng, Jinliang, Zhang, Wencong, Liu, Dongxiu, Zheng, Yinan, Niu, Haoyi, Yu, Junzhi, Zhan, Xianyuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918150485835776
author Wang, Zhihao
Li, Jianxiong
Zheng, Jinliang
Zhang, Wencong
Liu, Dongxiu
Zheng, Yinan
Niu, Haoyi
Yu, Junzhi
Zhan, Xianyuan
author_facet Wang, Zhihao
Li, Jianxiong
Zheng, Jinliang
Zhang, Wencong
Liu, Dongxiu
Zheng, Yinan
Niu, Haoyi
Yu, Junzhi
Zhan, Xianyuan
contents Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular solution. However, current approaches often combine these models in rigid, sequential structures: using VLMs primarily for high-level scene understanding and task planning, and VLAs merely as executors of lower-level actions, leading to ineffective collaboration and poor grounding challenges. In this paper, we propose an embodied agent framework, PhysiAgent, tailored to operate effectively in physical environments. By incorporating monitor, memory, self-reflection mechanisms, and lightweight off-the-shelf toolboxes, PhysiAgent offers an autonomous scaffolding framework to prompt VLMs to organize different components based on real-time proficiency feedback from VLAs to maximally exploit VLAs' capabilities. Experimental results demonstrate significant improvements in task-solving performance on complex real-world robotic tasks, showcasing effective self-regulation of VLMs, coherent tool collaboration, and adaptive evolution of the framework during execution. PhysiAgent makes practical and pioneering efforts to integrate VLMs and VLAs, effectively grounding embodied agent frameworks in real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24524
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhysiAgent: An Embodied Agent Framework in Physical World
Wang, Zhihao
Li, Jianxiong
Zheng, Jinliang
Zhang, Wencong
Liu, Dongxiu
Zheng, Yinan
Niu, Haoyi
Yu, Junzhi
Zhan, Xianyuan
Robotics
Artificial Intelligence
Systems and Control
Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations. To address this, integrating generalized Vision-Language Models (VLMs) as assistants to VLAs has emerged as a popular solution. However, current approaches often combine these models in rigid, sequential structures: using VLMs primarily for high-level scene understanding and task planning, and VLAs merely as executors of lower-level actions, leading to ineffective collaboration and poor grounding challenges. In this paper, we propose an embodied agent framework, PhysiAgent, tailored to operate effectively in physical environments. By incorporating monitor, memory, self-reflection mechanisms, and lightweight off-the-shelf toolboxes, PhysiAgent offers an autonomous scaffolding framework to prompt VLMs to organize different components based on real-time proficiency feedback from VLAs to maximally exploit VLAs' capabilities. Experimental results demonstrate significant improvements in task-solving performance on complex real-world robotic tasks, showcasing effective self-regulation of VLMs, coherent tool collaboration, and adaptive evolution of the framework during execution. PhysiAgent makes practical and pioneering efforts to integrate VLMs and VLAs, effectively grounding embodied agent frameworks in real-world settings.
title PhysiAgent: An Embodied Agent Framework in Physical World
topic Robotics
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2509.24524