EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Cheng, Han, Peixuan, Luo, Qinyu, He, Bingxiang, Chen, Xiusi, Zhang, Yuji, Du, Hongyi, Yao, Jiarui, Yang, Xiaocheng, Zhang, Denghui, Li, Yunzhu, Ji, Heng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908377326551040
author Qian, Cheng
Han, Peixuan
Luo, Qinyu
He, Bingxiang
Chen, Xiusi
Zhang, Yuji
Du, Hongyi
Yao, Jiarui
Yang, Xiaocheng
Zhang, Denghui
Li, Yunzhu
Ji, Heng
author_facet Qian, Cheng
Han, Peixuan
Luo, Qinyu
He, Bingxiang
Chen, Xiusi
Zhang, Yuji
Du, Hongyi
Yao, Jiarui
Yang, Xiaocheng
Zhang, Denghui
Li, Yunzhu
Ji, Heng
contents Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative adaptation in unfamiliar environments. To address this, we introduce EscapeBench, a benchmark suite of room escape game environments designed to challenge agents with creative reasoning, unconventional tool use, and iterative problem-solving to uncover implicit goals. Our results show that current LM models, despite employing working memory and Chain-of-Thought reasoning, achieve only 15% average progress without hints, highlighting their limitations in creativity. To bridge this gap, we propose EscapeAgent, a framework designed to enhance creative reasoning through Foresight (innovative tool use) and Reflection (identifying unsolved tasks). Experiments show that EscapeAgent can execute action chains over 1,000 steps while maintaining logical coherence. It navigates and completes games with up to 40% fewer steps and hints, performs robustly across difficulty levels, and achieves higher action success rates with more efficient and innovative puzzle-solving strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
Qian, Cheng
Han, Peixuan
Luo, Qinyu
He, Bingxiang
Chen, Xiusi
Zhang, Yuji
Du, Hongyi
Yao, Jiarui
Yang, Xiaocheng
Zhang, Denghui
Li, Yunzhu
Ji, Heng
Computation and Language
Artificial Intelligence
Machine Learning
Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative adaptation in unfamiliar environments. To address this, we introduce EscapeBench, a benchmark suite of room escape game environments designed to challenge agents with creative reasoning, unconventional tool use, and iterative problem-solving to uncover implicit goals. Our results show that current LM models, despite employing working memory and Chain-of-Thought reasoning, achieve only 15% average progress without hints, highlighting their limitations in creativity. To bridge this gap, we propose EscapeAgent, a framework designed to enhance creative reasoning through Foresight (innovative tool use) and Reflection (identifying unsolved tasks). Experiments show that EscapeAgent can execute action chains over 1,000 steps while maintaining logical coherence. It navigates and completes games with up to 40% fewer steps and hints, performs robustly across difficulty levels, and achieves higher action success rates with more efficient and innovative puzzle-solving strategies.
title EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.13549