EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Davies, Travis, Huang, Yiqi, Gladstone, Alexi, Liu, Yunxin, Chen, Xiang, Ji, Heng, Liu, Huxian, Hu, Luhui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917052817604608
author Davies, Travis
Huang, Yiqi
Gladstone, Alexi
Liu, Yunxin
Chen, Xiang
Ji, Heng
Liu, Huxian
Hu, Luhui
author_facet Davies, Travis
Huang, Yiqi
Gladstone, Alexi
Liu, Yunxin
Chen, Xiang
Ji, Heng
Liu, Huxian
Hu, Luhui
contents Implicit policies parameterized by generative models, such as Diffusion Policy, have become the standard for policy learning and Vision-Language-Action (VLA) models in robotics. However, these approaches often suffer from high computational cost, exposure bias, and unstable inference dynamics, which lead to divergence under distribution shifts. Energy-Based Models (EBMs) address these issues by learning energy landscapes end-to-end and modeling equilibrium dynamics, offering improved robustness and reduced exposure bias. Yet, policies parameterized by EBMs have historically struggled to scale effectively. Recent work on Energy-Based Transformers (EBTs) demonstrates the scalability of EBMs to high-dimensional spaces, but their potential for solving core challenges in physically embodied models remains underexplored. We introduce a new energy-based architecture, EBT-Policy, that solves core issues in robotic and real-world settings. Across simulated and real-world tasks, EBT-Policy consistently outperforms diffusion-based policies, while requiring less training and inference computation. Remarkably, on some tasks it converges within just two inference steps, a 50x reduction compared to Diffusion Policy's 100. Moreover, EBT-Policy exhibits emergent capabilities not seen in prior models, such as zero-shot recovery from failed action sequences using only behavior cloning and without explicit retry training. By leveraging its scalar energy for uncertainty-aware inference and dynamic compute allocation, EBT-Policy offers a promising path toward robust, generalizable robot behavior under distribution shifts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27545
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
Davies, Travis
Huang, Yiqi
Gladstone, Alexi
Liu, Yunxin
Chen, Xiang
Ji, Heng
Liu, Huxian
Hu, Luhui
Robotics
Artificial Intelligence
Implicit policies parameterized by generative models, such as Diffusion Policy, have become the standard for policy learning and Vision-Language-Action (VLA) models in robotics. However, these approaches often suffer from high computational cost, exposure bias, and unstable inference dynamics, which lead to divergence under distribution shifts. Energy-Based Models (EBMs) address these issues by learning energy landscapes end-to-end and modeling equilibrium dynamics, offering improved robustness and reduced exposure bias. Yet, policies parameterized by EBMs have historically struggled to scale effectively. Recent work on Energy-Based Transformers (EBTs) demonstrates the scalability of EBMs to high-dimensional spaces, but their potential for solving core challenges in physically embodied models remains underexplored. We introduce a new energy-based architecture, EBT-Policy, that solves core issues in robotic and real-world settings. Across simulated and real-world tasks, EBT-Policy consistently outperforms diffusion-based policies, while requiring less training and inference computation. Remarkably, on some tasks it converges within just two inference steps, a 50x reduction compared to Diffusion Policy's 100. Moreover, EBT-Policy exhibits emergent capabilities not seen in prior models, such as zero-shot recovery from failed action sequences using only behavior cloning and without explicit retry training. By leveraging its scalar energy for uncertainty-aware inference and dynamic compute allocation, EBT-Policy offers a promising path toward robust, generalizable robot behavior under distribution shifts.
title EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2510.27545