TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hudi, Frederikus, Winata, Genta Indra, Zhang, Ruochen, Aji, Alham Fikri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912245777170432
author Hudi, Frederikus
Winata, Genta Indra
Zhang, Ruochen
Aji, Alham Fikri
author_facet Hudi, Frederikus
Winata, Genta Indra
Zhang, Ruochen
Aji, Alham Fikri
contents Reasoning is a fundamental capability of large language models (LLMs), enabling them to comprehend, analyze, and solve complex problems. In this paper, we introduce TextGames, an innovative benchmark specifically crafted to assess LLMs through demanding text-based games that require advanced skills in pattern recognition, spatial awareness, arithmetic, and logical reasoning. Our analysis probes LLMs' performance in both single-turn and multi-turn reasoning, and their abilities in leveraging feedback to correct subsequent answers through self-reflection. Our findings reveal that, although LLMs exhibit proficiency in addressing most easy and medium-level problems, they face significant challenges with more difficult tasks. In contrast, humans are capable of solving all tasks when given sufficient time. Moreover, we observe that LLMs show improved performance in multi-turn predictions through self-reflection, yet they still struggle with sequencing, counting, and following complex rules consistently. Additionally, models optimized for reasoning outperform pre-trained LLMs that prioritize instruction following, highlighting the crucial role of reasoning skills in addressing highly complex problems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
Hudi, Frederikus
Winata, Genta Indra
Zhang, Ruochen
Aji, Alham Fikri
Computation and Language
Artificial Intelligence
Reasoning is a fundamental capability of large language models (LLMs), enabling them to comprehend, analyze, and solve complex problems. In this paper, we introduce TextGames, an innovative benchmark specifically crafted to assess LLMs through demanding text-based games that require advanced skills in pattern recognition, spatial awareness, arithmetic, and logical reasoning. Our analysis probes LLMs' performance in both single-turn and multi-turn reasoning, and their abilities in leveraging feedback to correct subsequent answers through self-reflection. Our findings reveal that, although LLMs exhibit proficiency in addressing most easy and medium-level problems, they face significant challenges with more difficult tasks. In contrast, humans are capable of solving all tasks when given sufficient time. Moreover, we observe that LLMs show improved performance in multi-turn predictions through self-reflection, yet they still struggle with sequencing, counting, and following complex rules consistently. Additionally, models optimized for reasoning outperform pre-trained LLMs that prioritize instruction following, highlighting the crucial role of reasoning skills in addressing highly complex problems.
title TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.18431