Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sokota, Samuel, Vinitsky, Eugene, Hu, Hengyuan, Kolter, J. Zico, Farina, Gabriele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912698962280448
author Sokota, Samuel
Vinitsky, Eugene
Hu, Hengyuan
Kolter, J. Zico
Farina, Gabriele
author_facet Sokota, Samuel
Vinitsky, Eugene
Hu, Hengyuan
Kolter, J. Zico
Farina, Gabriele
contents Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Stratego -- a board wargame exemplifying the challenge of strategic decision making under massive amounts of hidden information -- stands apart as a case where such efforts failed to produce performance at the level of top humans. This work establishes a step change in both performance and cost for Stratego, showing that it is now possible not only to reach the level of top humans, but to achieve vastly superhuman level -- and that doing so requires not an industrial budget, but merely a few thousand dollars. We achieved this result by developing general approaches for self-play reinforcement learning and test-time search under imperfect information.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
Sokota, Samuel
Vinitsky, Eugene
Hu, Hengyuan
Kolter, J. Zico
Farina, Gabriele
Machine Learning
Artificial Intelligence
Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Stratego -- a board wargame exemplifying the challenge of strategic decision making under massive amounts of hidden information -- stands apart as a case where such efforts failed to produce performance at the level of top humans. This work establishes a step change in both performance and cost for Stratego, showing that it is now possible not only to reach the level of top humans, but to achieve vastly superhuman level -- and that doing so requires not an industrial budget, but merely a few thousand dollars. We achieved this result by developing general approaches for self-play reinforcement learning and test-time search under imperfect information.
title Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.07312