Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jianzong, Zhao, Botao, He, Yayun, Peng, Junqing, Zhang, Xulong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908985528942592
author Wang, Jianzong
Zhao, Botao
He, Yayun
Peng, Junqing
Zhang, Xulong
author_facet Wang, Jianzong
Zhao, Botao
He, Yayun
Peng, Junqing
Zhang, Xulong
contents Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive training requirements, difficulties in cross-task generalization, and lack of interpretability. Prompt learning offers new opportunities for self-evolving robots without extensive training, but simply reflecting on past experiences. However, extracting meaningful insights from task successes and failures remains a challenge. To this end, we propose the evolvable embodied agent (EEAgent) framework, which leverages large vision-language models (VLMs) for better environmental interpretation and policy planning. To enhance reflection on past experiences, we propose a long short-term reflective optimization (LSTRO) mechanism that dynamically refines prompts based on both past experiences and newly learned lessons, facilitating continuous self-evolution, thereby enhancing overall task success rates. Evaluations on six VIMA-Bench tasks reveal that our approach sets a new state-of-the-art, notably outperforming baselines in complex scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13533
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
Wang, Jianzong
Zhao, Botao
He, Yayun
Peng, Junqing
Zhang, Xulong
Robotics
Computer Vision and Pattern Recognition
Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face limitations such as extensive training requirements, difficulties in cross-task generalization, and lack of interpretability. Prompt learning offers new opportunities for self-evolving robots without extensive training, but simply reflecting on past experiences. However, extracting meaningful insights from task successes and failures remains a challenge. To this end, we propose the evolvable embodied agent (EEAgent) framework, which leverages large vision-language models (VLMs) for better environmental interpretation and policy planning. To enhance reflection on past experiences, we propose a long short-term reflective optimization (LSTRO) mechanism that dynamically refines prompts based on both past experiences and newly learned lessons, facilitating continuous self-evolution, thereby enhancing overall task success rates. Evaluations on six VIMA-Bench tasks reveal that our approach sets a new state-of-the-art, notably outperforming baselines in complex scenarios.
title Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.13533