ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hanyang, Zhao, Mark, Yang, Rui, Ma, Qinwei, Yang, Ke, Yao, Jiarui, Wang, Kangrui, Bai, Hao, Wang, Zhenhailong, Pan, Rui, Zhang, Mengchao, Barreiros, Jose, Onol, Aykut, Zhai, ChengXiang, Ji, Heng, Li, Manling, Zhang, Huan, Zhang, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912647034699776
author Chen, Hanyang
Zhao, Mark
Yang, Rui
Ma, Qinwei
Yang, Ke
Yao, Jiarui
Wang, Kangrui
Bai, Hao
Wang, Zhenhailong
Pan, Rui
Zhang, Mengchao
Barreiros, Jose
Onol, Aykut
Zhai, ChengXiang
Ji, Heng
Li, Manling
Zhang, Huan
Zhang, Tong
author_facet Chen, Hanyang
Zhao, Mark
Yang, Rui
Ma, Qinwei
Yang, Ke
Yao, Jiarui
Wang, Kangrui
Bai, Hao
Wang, Zhenhailong
Pan, Rui
Zhang, Mengchao
Barreiros, Jose
Onol, Aykut
Zhai, ChengXiang
Ji, Heng
Li, Manling
Zhang, Huan
Zhang, Tong
contents Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However, top-performing systems rely on large-scale models that are costly to deploy, while smaller VLMs lack the necessary knowledge and skills to succeed. To bridge this gap, we present \textit{Embodied Reasoning Agent (ERA)}, a two-stage framework that integrates prior knowledge learning and online reinforcement learning (RL). The first stage, \textit{Embodied Prior Learning}, distills foundational knowledge from three types of data: (1) Trajectory-Augmented Priors, which enrich existing trajectory data with structured reasoning generated by stronger models; (2) Environment-Anchored Priors, which provide in-environment knowledge and grounding supervision; and (3) External Knowledge Priors, which transfer general knowledge from out-of-environment datasets. In the second stage, we develop an online RL pipeline that builds on these priors to further enhance agent performance. To overcome the inherent challenges in agent RL, including long horizons, sparse rewards, and training instability, we introduce three key designs: self-summarization for context management, dense reward shaping, and turn-level policy optimization. Extensive experiments on both high-level planning (EB-ALFRED) and low-level control (EB-Manipulation) tasks demonstrate that ERA-3B surpasses both prompting-based large models and previous training-based baselines. Specifically, it achieves overall improvements of 8.4\% on EB-ALFRED and 19.4\% on EB-Manipulation over GPT-4o, and exhibits strong generalization to unseen tasks. Overall, ERA offers a practical path toward scalable embodied intelligence, providing methodological insights for future embodied AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12693
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
Chen, Hanyang
Zhao, Mark
Yang, Rui
Ma, Qinwei
Yang, Ke
Yao, Jiarui
Wang, Kangrui
Bai, Hao
Wang, Zhenhailong
Pan, Rui
Zhang, Mengchao
Barreiros, Jose
Onol, Aykut
Zhai, ChengXiang
Ji, Heng
Li, Manling
Zhang, Huan
Zhang, Tong
Artificial Intelligence
Recent advances in embodied AI highlight the potential of vision language models (VLMs) as agents capable of perception, reasoning, and interaction in complex environments. However, top-performing systems rely on large-scale models that are costly to deploy, while smaller VLMs lack the necessary knowledge and skills to succeed. To bridge this gap, we present \textit{Embodied Reasoning Agent (ERA)}, a two-stage framework that integrates prior knowledge learning and online reinforcement learning (RL). The first stage, \textit{Embodied Prior Learning}, distills foundational knowledge from three types of data: (1) Trajectory-Augmented Priors, which enrich existing trajectory data with structured reasoning generated by stronger models; (2) Environment-Anchored Priors, which provide in-environment knowledge and grounding supervision; and (3) External Knowledge Priors, which transfer general knowledge from out-of-environment datasets. In the second stage, we develop an online RL pipeline that builds on these priors to further enhance agent performance. To overcome the inherent challenges in agent RL, including long horizons, sparse rewards, and training instability, we introduce three key designs: self-summarization for context management, dense reward shaping, and turn-level policy optimization. Extensive experiments on both high-level planning (EB-ALFRED) and low-level control (EB-Manipulation) tasks demonstrate that ERA-3B surpasses both prompting-based large models and previous training-based baselines. Specifically, it achieves overall improvements of 8.4\% on EB-ALFRED and 19.4\% on EB-Manipulation over GPT-4o, and exhibits strong generalization to unseen tasks. Overall, ERA offers a practical path toward scalable embodied intelligence, providing methodological insights for future embodied AI systems.
title ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2510.12693