ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jincheng, He, Sijun, Wu, Jingjing, Wang, Xiangsen, Chen, Yang, Kuang, Zhaoqi, Bao, Siqi, Yao, Yuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908988511092736
author Liu, Jincheng
He, Sijun
Wu, Jingjing
Wang, Xiangsen
Chen, Yang
Kuang, Zhaoqi
Bao, Siqi
Yao, Yuan
author_facet Liu, Jincheng
He, Sijun
Wu, Jingjing
Wang, Xiangsen
Chen, Yang
Kuang, Zhaoqi
Bao, Siqi
Yao, Yuan
contents Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess genuine strategic reasoning, or do they primarily excel at pattern recognition? To address this, we present ChessArena, a chess-based testbed for evaluating LLMs. Chess demands strategic reasoning, precise rule adherence, and the ability to track complex game states. ChessArena is a competitive framework where LLMs play against each other under four play modes. We evaluate 13 LLMs across over 800 games, testing basic understanding, move selection, and puzzle solving. Results reveal significant shortcomings: no model beats Maia-1100 (human amateur level), and some lose to random play. We also present a strong baseline: our fine-tuned Qwen3-8B substantially improves performance, approaching much larger state-of-the-art reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24239
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
Liu, Jincheng
He, Sijun
Wu, Jingjing
Wang, Xiangsen
Chen, Yang
Kuang, Zhaoqi
Bao, Siqi
Yao, Yuan
Machine Learning
Artificial Intelligence
Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess genuine strategic reasoning, or do they primarily excel at pattern recognition? To address this, we present ChessArena, a chess-based testbed for evaluating LLMs. Chess demands strategic reasoning, precise rule adherence, and the ability to track complex game states. ChessArena is a competitive framework where LLMs play against each other under four play modes. We evaluate 13 LLMs across over 800 games, testing basic understanding, move selection, and puzzle solving. Results reveal significant shortcomings: no model beats Maia-1100 (human amateur level), and some lose to random play. We also present a strong baseline: our fine-tuned Qwen3-8B substantially improves performance, approaching much larger state-of-the-art reasoning models.
title ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.24239