How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Minhua, Dai, Enyan, Liu, Hui, Tang, Xianfeng, Yan, Yuliang, Dai, Zhenwei, Zeng, Jingying, Zhang, Zhiwei, Wang, Fali, Gao, Hongcheng, Luo, Chen, Zhang, Xiang, He, Qi, Wang, Suhang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911412985528320
author Lin, Minhua
Dai, Enyan
Liu, Hui
Tang, Xianfeng
Yan, Yuliang
Dai, Zhenwei
Zeng, Jingying
Zhang, Zhiwei
Wang, Fali
Gao, Hongcheng
Luo, Chen
Zhang, Xiang
He, Qi
Wang, Suhang
author_facet Lin, Minhua
Dai, Enyan
Liu, Hui
Tang, Xianfeng
Yan, Yuliang
Dai, Zhenwei
Zeng, Jingying
Zhang, Zhiwei
Wang, Fali
Gao, Hongcheng
Luo, Chen
Zhang, Xiang
He, Qi
Wang, Suhang
contents As Large Language Models (LLMs) are increasingly applied in high-stakes domains, their ability to reason strategically under uncertainty becomes critical. Poker provides a rigorous testbed, requiring not only strong actions but also principled, game-theoretic reasoning. In this paper, we conduct a systematic study of LLMs in multiple realistic poker tasks, evaluating both gameplay outcomes and reasoning traces. Our analysis reveals LLMs fail to compete against traditional algorithms and identifies three recurring flaws: reliance on heuristics, factual misunderstandings, and a "knowing-doing" gap where actions diverge from reasoning. An initial attempt with behavior cloning and step-level reinforcement learning improves reasoning style but remains insufficient for accurate game-theoretic play. Motivated by these limitations, we propose ToolPoker, a tool-integrated reasoning framework that combines external solvers for GTO-consistent actions with more precise professional-style explanations. Experiments demonstrate that ToolPoker achieves state-of-the-art gameplay while producing reasoning traces that closely reflect game-theoretic principles.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00528
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
Lin, Minhua
Dai, Enyan
Liu, Hui
Tang, Xianfeng
Yan, Yuliang
Dai, Zhenwei
Zeng, Jingying
Zhang, Zhiwei
Wang, Fali
Gao, Hongcheng
Luo, Chen
Zhang, Xiang
He, Qi
Wang, Suhang
Artificial Intelligence
As Large Language Models (LLMs) are increasingly applied in high-stakes domains, their ability to reason strategically under uncertainty becomes critical. Poker provides a rigorous testbed, requiring not only strong actions but also principled, game-theoretic reasoning. In this paper, we conduct a systematic study of LLMs in multiple realistic poker tasks, evaluating both gameplay outcomes and reasoning traces. Our analysis reveals LLMs fail to compete against traditional algorithms and identifies three recurring flaws: reliance on heuristics, factual misunderstandings, and a "knowing-doing" gap where actions diverge from reasoning. An initial attempt with behavior cloning and step-level reinforcement learning improves reasoning style but remains insufficient for accurate game-theoretic play. Motivated by these limitations, we propose ToolPoker, a tool-integrated reasoning framework that combines external solvers for GTO-consistent actions with more precise professional-style explanations. Experiments demonstrate that ToolPoker achieves state-of-the-art gameplay while producing reasoning traces that closely reflect game-theoretic principles.
title How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
topic Artificial Intelligence
url https://arxiv.org/abs/2602.00528