Automated Reward Design for Gran Turismo
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914133648080896 |
|---|---|
| author | Ma, Michel Seno, Takuma Subramanian, Kaushik Wurman, Peter R. Stone, Peter Sherstan, Craig |
| author_facet | Ma, Michel Seno, Takuma Subramanian, Kaushik Wurman, Peter R. Stone, Peter Sherstan, Craig |
| contents | When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the agent as reward or punishment for its actions. However, mapping desired behaviors to reward functions can be a difficult process, especially in complex environments such as autonomous racing. In this paper, we demonstrate how current foundation models can effectively search over a space of reward functions to produce desirable RL agents for the Gran Turismo 7 racing game, given only text-based instructions. Through a combination of LLM-based reward generation, VLM preference-based evaluation, and human feedback we demonstrate how our system can be used to produce racing agents competitive with GT Sophy, a champion-level RL racing agent, as well as generate novel behaviors, paving the way for practical automated reward design in real world applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_02094 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Automated Reward Design for Gran Turismo Ma, Michel Seno, Takuma Subramanian, Kaushik Wurman, Peter R. Stone, Peter Sherstan, Craig Artificial Intelligence When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the agent as reward or punishment for its actions. However, mapping desired behaviors to reward functions can be a difficult process, especially in complex environments such as autonomous racing. In this paper, we demonstrate how current foundation models can effectively search over a space of reward functions to produce desirable RL agents for the Gran Turismo 7 racing game, given only text-based instructions. Through a combination of LLM-based reward generation, VLM preference-based evaluation, and human feedback we demonstrate how our system can be used to produce racing agents competitive with GT Sophy, a champion-level RL racing agent, as well as generate novel behaviors, paving the way for practical automated reward design in real world applications. |
| title | Automated Reward Design for Gran Turismo |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2511.02094 |