Explore the Reasoning Capability of LLMs in the Chess Testbed

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Shu, Ji, Lei, Wang, Renxi, Zhao, Wenxiao, Liu, Haokun, Hou, Yifan, Wu, Ying Nian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915176209448960
author Wang, Shu
Ji, Lei
Wang, Renxi
Zhao, Wenxiao
Liu, Haokun
Hou, Yifan
Wu, Ying Nian
author_facet Wang, Shu
Ji, Lei
Wang, Renxi
Zhao, Wenxiao
Liu, Haokun
Hou, Yifan
Wu, Ying Nian
contents Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing chess. Based on the observation that expert chess players employ a dual approach combining long-term strategic play with short-term tactical play along with language explanation, we propose improving the reasoning capability of large language models in chess by integrating annotated strategy and tactic. Specifically, we collect a dataset named MATE, which consists of 1 million chess positions with candidate moves annotated by chess experts for strategy and tactics. We finetune the LLaMA-3-8B model and compare it against state-of-the-art commercial language models in the task of selecting better chess moves. Our experiments show that our models perform better than GPT, Claude, and Gemini models. We find that language explanations can enhance the reasoning capability of large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06655
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explore the Reasoning Capability of LLMs in the Chess Testbed
Wang, Shu
Ji, Lei
Wang, Renxi
Zhao, Wenxiao
Liu, Haokun
Hou, Yifan
Wu, Ying Nian
Computation and Language
Artificial Intelligence
Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing chess. Based on the observation that expert chess players employ a dual approach combining long-term strategic play with short-term tactical play along with language explanation, we propose improving the reasoning capability of large language models in chess by integrating annotated strategy and tactic. Specifically, we collect a dataset named MATE, which consists of 1 million chess positions with candidate moves annotated by chess experts for strategy and tactics. We finetune the LLaMA-3-8B model and compare it against state-of-the-art commercial language models in the task of selecting better chess moves. Our experiments show that our models perform better than GPT, Claude, and Gemini models. We find that language explanations can enhance the reasoning capability of large language models.
title Explore the Reasoning Capability of LLMs in the Chess Testbed
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.06655