Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914472960983040 |
|---|---|
| author | Park, Dongmin Kim, Minkyu Choi, Beongjun Kim, Junhyuck Lee, Keon Lee, Jonghyun Park, Inkyu Lee, Byeong-Uk Hwang, Jaeyoung Ahn, Jaewoo Mahabaleshwarkar, Ameya S. Kartal, Bilal Biswas, Pritam Suhara, Yoshi Lee, Kangwook Cho, Jaewoong |
| author_facet | Park, Dongmin Kim, Minkyu Choi, Beongjun Kim, Junhyuck Lee, Keon Lee, Jonghyun Park, Inkyu Lee, Byeong-Uk Hwang, Jaeyoung Ahn, Jaewoo Mahabaleshwarkar, Ameya S. Kartal, Bilal Biswas, Pritam Suhara, Yoshi Lee, Kangwook Cho, Jaewoong |
| contents | Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories covering multiple genres, turning general LLMs into effective game agents. Orak offers a united evaluation framework, including game leaderboards, LLM battle arenas, and \fix{ablation studies} of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_03610 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games Park, Dongmin Kim, Minkyu Choi, Beongjun Kim, Junhyuck Lee, Keon Lee, Jonghyun Park, Inkyu Lee, Byeong-Uk Hwang, Jaeyoung Ahn, Jaewoo Mahabaleshwarkar, Ameya S. Kartal, Bilal Biswas, Pritam Suhara, Yoshi Lee, Kangwook Cho, Jaewoong Artificial Intelligence Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories covering multiple genres, turning general LLMs into effective game agents. Orak offers a united evaluation framework, including game leaderboards, LLM battle arenas, and \fix{ablation studies} of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak. |
| title | Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2506.03610 |