Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Dongmin, Kim, Minkyu, Choi, Beongjun, Kim, Junhyuck, Lee, Keon, Lee, Jonghyun, Park, Inkyu, Lee, Byeong-Uk, Hwang, Jaeyoung, Ahn, Jaewoo, Mahabaleshwarkar, Ameya S., Kartal, Bilal, Biswas, Pritam, Suhara, Yoshi, Lee, Kangwook, Cho, Jaewoong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914472960983040
author Park, Dongmin
Kim, Minkyu
Choi, Beongjun
Kim, Junhyuck
Lee, Keon
Lee, Jonghyun
Park, Inkyu
Lee, Byeong-Uk
Hwang, Jaeyoung
Ahn, Jaewoo
Mahabaleshwarkar, Ameya S.
Kartal, Bilal
Biswas, Pritam
Suhara, Yoshi
Lee, Kangwook
Cho, Jaewoong
author_facet Park, Dongmin
Kim, Minkyu
Choi, Beongjun
Kim, Junhyuck
Lee, Keon
Lee, Jonghyun
Park, Inkyu
Lee, Byeong-Uk
Hwang, Jaeyoung
Ahn, Jaewoo
Mahabaleshwarkar, Ameya S.
Kartal, Bilal
Biswas, Pritam
Suhara, Yoshi
Lee, Kangwook
Cho, Jaewoong
contents Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories covering multiple genres, turning general LLMs into effective game agents. Orak offers a united evaluation framework, including game leaderboards, LLM battle arenas, and \fix{ablation studies} of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Park, Dongmin
Kim, Minkyu
Choi, Beongjun
Kim, Junhyuck
Lee, Keon
Lee, Jonghyun
Park, Inkyu
Lee, Byeong-Uk
Hwang, Jaeyoung
Ahn, Jaewoo
Mahabaleshwarkar, Ameya S.
Kartal, Bilal
Biswas, Pritam
Suhara, Yoshi
Lee, Kangwook
Cho, Jaewoong
Artificial Intelligence
Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories covering multiple genres, turning general LLMs into effective game agents. Orak offers a united evaluation framework, including game leaderboards, LLM battle arenas, and \fix{ablation studies} of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code and datasets are available at https://github.com/krafton-ai/Orak and https://huggingface.co/datasets/KRAFTON/Orak.
title Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
topic Artificial Intelligence
url https://arxiv.org/abs/2506.03610