EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Zixing, Liu, Genjia, Zhang, Yuanshuo, Liu, Qipeng, Wen, Chuan, Zhang, Shanghang, Lian, Wenzhao, Chen, Siheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917231433089024
author Lei, Zixing
Liu, Genjia
Zhang, Yuanshuo
Liu, Qipeng
Wen, Chuan
Zhang, Shanghang
Lian, Wenzhao
Chen, Siheng
author_facet Lei, Zixing
Liu, Genjia
Zhang, Yuanshuo
Liu, Qipeng
Wen, Chuan
Zhang, Shanghang
Lian, Wenzhao
Chen, Siheng
contents The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this scaling capability remains severely bottlenecked by a reliance on labor-intensive manual oversight from intricate reward shaping to hyperparameter tuning across heterogeneous backends. Inspired by LLMs' success in software automation and science discovery, we introduce \textsc{EmboCoach-Bench}, a benchmark evaluating the capacity of LLM agents to autonomously engineer embodied policies. Spanning 32 expert-curated RL and IL tasks, our framework posits executable code as the universal interface. We move beyond static generation to assess a dynamic closed-loop workflow, where agents leverage environment feedback to iteratively draft, debug, and optimize solutions, spanning improvements from physics-informed reward design to policy architectures such as diffusion policies. Extensive evaluations yield three critical insights: (1) autonomous agents can qualitatively surpass human-engineered baselines by 26.5\% in average success rate; (2) agentic workflow with environment feedback effectively strengthens policy development and substantially narrows the performance gap between open-source and proprietary models; and (3) agents exhibit self-correction capabilities for pathological engineering cases, successfully resurrecting task performance from near-total failures through iterative simulation-in-the-loop debugging. Ultimately, this work establishes a foundation for self-evolving embodied intelligence, accelerating the paradigm shift from labor-intensive manual tuning to scalable, autonomous engineering in embodied AI field.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21570
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
Lei, Zixing
Liu, Genjia
Zhang, Yuanshuo
Liu, Qipeng
Wen, Chuan
Zhang, Shanghang
Lian, Wenzhao
Chen, Siheng
Artificial Intelligence
Robotics
The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this scaling capability remains severely bottlenecked by a reliance on labor-intensive manual oversight from intricate reward shaping to hyperparameter tuning across heterogeneous backends. Inspired by LLMs' success in software automation and science discovery, we introduce \textsc{EmboCoach-Bench}, a benchmark evaluating the capacity of LLM agents to autonomously engineer embodied policies. Spanning 32 expert-curated RL and IL tasks, our framework posits executable code as the universal interface. We move beyond static generation to assess a dynamic closed-loop workflow, where agents leverage environment feedback to iteratively draft, debug, and optimize solutions, spanning improvements from physics-informed reward design to policy architectures such as diffusion policies. Extensive evaluations yield three critical insights: (1) autonomous agents can qualitatively surpass human-engineered baselines by 26.5\% in average success rate; (2) agentic workflow with environment feedback effectively strengthens policy development and substantially narrows the performance gap between open-source and proprietary models; and (3) agents exhibit self-correction capabilities for pathological engineering cases, successfully resurrecting task performance from near-total failures through iterative simulation-in-the-loop debugging. Ultimately, this work establishes a foundation for self-evolving embodied intelligence, accelerating the paradigm shift from labor-intensive manual tuning to scalable, autonomous engineering in embodied AI field.
title EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2601.21570