Automated Reward Design for Gran Turismo

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Michel, Seno, Takuma, Subramanian, Kaushik, Wurman, Peter R., Stone, Peter, Sherstan, Craig
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914133648080896
author Ma, Michel
Seno, Takuma
Subramanian, Kaushik
Wurman, Peter R.
Stone, Peter
Sherstan, Craig
author_facet Ma, Michel
Seno, Takuma
Subramanian, Kaushik
Wurman, Peter R.
Stone, Peter
Sherstan, Craig
contents When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the agent as reward or punishment for its actions. However, mapping desired behaviors to reward functions can be a difficult process, especially in complex environments such as autonomous racing. In this paper, we demonstrate how current foundation models can effectively search over a space of reward functions to produce desirable RL agents for the Gran Turismo 7 racing game, given only text-based instructions. Through a combination of LLM-based reward generation, VLM preference-based evaluation, and human feedback we demonstrate how our system can be used to produce racing agents competitive with GT Sophy, a champion-level RL racing agent, as well as generate novel behaviors, paving the way for practical automated reward design in real world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02094
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Reward Design for Gran Turismo
Ma, Michel
Seno, Takuma
Subramanian, Kaushik
Wurman, Peter R.
Stone, Peter
Sherstan, Craig
Artificial Intelligence
When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the agent as reward or punishment for its actions. However, mapping desired behaviors to reward functions can be a difficult process, especially in complex environments such as autonomous racing. In this paper, we demonstrate how current foundation models can effectively search over a space of reward functions to produce desirable RL agents for the Gran Turismo 7 racing game, given only text-based instructions. Through a combination of LLM-based reward generation, VLM preference-based evaluation, and human feedback we demonstrate how our system can be used to produce racing agents competitive with GT Sophy, a champion-level RL racing agent, as well as generate novel behaviors, paving the way for practical automated reward design in real world applications.
title Automated Reward Design for Gran Turismo
topic Artificial Intelligence
url https://arxiv.org/abs/2511.02094