Pokemon Red via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pleines, Marco, Addis, Daniel, Rubinstein, David, Zimmer, Frank, Preuss, Mike, Whidden, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912269248495616
author Pleines, Marco
Addis, Daniel
Rubinstein, David
Zimmer, Frank
Preuss, Mike
Whidden, Peter
author_facet Pleines, Marco
Addis, Daniel
Rubinstein, David
Zimmer, Frank
Preuss, Mike
Whidden, Peter
contents Pokémon Red, a classic Game Boy JRPG, presents significant challenges as a testbed for agents, including multi-tasking, long horizons of tens of thousands of steps, hard exploration, and a vast array of potential policies. We introduce a simplistic environment and a Deep Reinforcement Learning (DRL) training methodology, demonstrating a baseline agent that completes an initial segment of the game up to completing Cerulean City. Our experiments include various ablations that reveal vulnerabilities in reward shaping, where agents exploit specific reward signals. We also discuss limitations and argue that games like Pokémon hold strong potential for future research on Large Language Model agents, hierarchical training algorithms, and advanced exploration methods. Source Code: https://github.com/MarcoMeter/neroRL/tree/poke_red
format Preprint
id arxiv_https___arxiv_org_abs_2502_19920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pokemon Red via Reinforcement Learning
Pleines, Marco
Addis, Daniel
Rubinstein, David
Zimmer, Frank
Preuss, Mike
Whidden, Peter
Machine Learning
Pokémon Red, a classic Game Boy JRPG, presents significant challenges as a testbed for agents, including multi-tasking, long horizons of tens of thousands of steps, hard exploration, and a vast array of potential policies. We introduce a simplistic environment and a Deep Reinforcement Learning (DRL) training methodology, demonstrating a baseline agent that completes an initial segment of the game up to completing Cerulean City. Our experiments include various ablations that reveal vulnerabilities in reward shaping, where agents exploit specific reward signals. We also discuss limitations and argue that games like Pokémon hold strong potential for future research on Large Language Model agents, hierarchical training algorithms, and advanced exploration methods. Source Code: https://github.com/MarcoMeter/neroRL/tree/poke_red
title Pokemon Red via Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2502.19920