LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhuoling, Xu, Xiaogang, Xu, Zhenhua, Lim, SerNam, Zhao, Hengshuang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912221120954368
author Li, Zhuoling
Xu, Xiaogang
Xu, Zhenhua
Lim, SerNam
Zhao, Hengshuang
author_facet Li, Zhuoling
Xu, Xiaogang
Xu, Zhenhua
Lim, SerNam
Zhao, Hengshuang
contents Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while demanding enormous computing resources. In this work, we combine their advantages while avoiding the drawbacks by conducting the proposed referee RL on our developed large auto-regressive model (LARM). Specifically, LARM is built upon a lightweight LLM (fewer than 5B parameters) and directly outputs the next action to execute rather than text. We mathematically reveal that classic RL feedbacks vanish in long-horizon embodied exploration and introduce a giant LLM based referee to handle this reward vanishment during training LARM. In this way, LARM learns to complete diverse open-world tasks without human intervention. Especially, LARM successfully harvests enchanted diamond equipment in Minecraft, which demands significantly longer decision-making chains than the highest achievements of prior best methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
Li, Zhuoling
Xu, Xiaogang
Xu, Zhenhua
Lim, SerNam
Zhao, Hengshuang
Computer Vision and Pattern Recognition
Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while demanding enormous computing resources. In this work, we combine their advantages while avoiding the drawbacks by conducting the proposed referee RL on our developed large auto-regressive model (LARM). Specifically, LARM is built upon a lightweight LLM (fewer than 5B parameters) and directly outputs the next action to execute rather than text. We mathematically reveal that classic RL feedbacks vanish in long-horizon embodied exploration and introduce a giant LLM based referee to handle this reward vanishment during training LARM. In this way, LARM learns to complete diverse open-world tasks without human intervention. Especially, LARM successfully harvests enchanted diamond equipment in Minecraft, which demands significantly longer decision-making chains than the highest achievements of prior best methods.
title LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.17424