FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Jaewoo, Kim, Junseo, Yun, Heeseung, Son, Jaehyeon, Park, Dongmin, Cho, Jaewoong, Kim, Gunhee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911210654400512
author Ahn, Jaewoo
Kim, Junseo
Yun, Heeseung
Son, Jaehyeon
Park, Dongmin
Cho, Jaewoong
Kim, Gunhee
author_facet Ahn, Jaewoo
Kim, Junseo
Yun, Heeseung
Son, Jaehyeon
Park, Dongmin
Cho, Jaewoong
Kim, Gunhee
contents GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions. Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines. To address this, we introduce FlashAdventure, a benchmark of 34 Flash-based adventure games designed to test full story arc completion and tackle the observation-behavior gap: the challenge of remembering and acting on earlier gameplay information. We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks. Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observation-behavior gap. Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01052
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
Ahn, Jaewoo
Kim, Junseo
Yun, Heeseung
Son, Jaehyeon
Park, Dongmin
Cho, Jaewoong
Kim, Gunhee
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions. Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines. To address this, we introduce FlashAdventure, a benchmark of 34 Flash-based adventure games designed to test full story arc completion and tackle the observation-behavior gap: the challenge of remembering and acting on earlier gameplay information. We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks. Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observation-behavior gap. Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide.
title FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.01052