GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sapora, Silvia, Hjelm, Devon, Toshev, Alexander, Attia, Omar, Mazoure, Bogdan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918311451688960
author Sapora, Silvia
Hjelm, Devon
Toshev, Alexander
Attia, Omar
Mazoure, Bogdan
author_facet Sapora, Silvia
Hjelm, Devon
Toshev, Alexander
Attia, Omar
Mazoure, Bogdan
contents Inverse Reinforcement Learning aims to recover reward models from expert demonstrations, but traditional methods yield black-box models that are difficult to interpret and debug. In this work, we introduce GRACE (Generating Rewards As CodE), a method for using Large Language Models within an evolutionary search to reverse-engineer an interpretable, code-based reward function directly from expert trajectories. The resulting reward function is executable code that can be inspected and verified. We empirically validate GRACE on the MuJoCo, BabyAI and AndroidWorld benchmarks, where it efficiently learns highly accurate rewards, even in complex, multi-task settings. Further, we demonstrate that the resulting reward leads to strong policies, compared to both competitive Imitation Learning and online RL approaches with ground-truth rewards. Finally, we show that GRACE is able to build complex reward APIs in multi-task setups.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
Sapora, Silvia
Hjelm, Devon
Toshev, Alexander
Attia, Omar
Mazoure, Bogdan
Machine Learning
Artificial Intelligence
Inverse Reinforcement Learning aims to recover reward models from expert demonstrations, but traditional methods yield black-box models that are difficult to interpret and debug. In this work, we introduce GRACE (Generating Rewards As CodE), a method for using Large Language Models within an evolutionary search to reverse-engineer an interpretable, code-based reward function directly from expert trajectories. The resulting reward function is executable code that can be inspected and verified. We empirically validate GRACE on the MuJoCo, BabyAI and AndroidWorld benchmarks, where it efficiently learns highly accurate rewards, even in complex, multi-task settings. Further, we demonstrate that the resulting reward leads to strong policies, compared to both competitive Imitation Learning and online RL approaches with ground-truth rewards. Finally, we show that GRACE is able to build complex reward APIs in multi-task setups.
title GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.02180