Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yikai, Rong, Ye, Yuan, Siyu, Chen, Jiangjie, Xie, Jian, Xiao, Yanghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915563103584256
author Zhang, Yikai
Rong, Ye
Yuan, Siyu
Chen, Jiangjie
Xie, Jian
Xiao, Yanghua
author_facet Zhang, Yikai
Rong, Ye
Yuan, Siyu
Chen, Jiangjie
Xie, Jian
Xiao, Yanghua
contents Existing language agents often encounter difficulties in dynamic adversarial games due to poor strategic reasoning. To mitigate this limitation, a promising approach is to allow agents to learn from game interactions automatically, without relying on costly expert-labeled data. Unlike static environments where agents receive fixed feedback or rewards, selecting appropriate opponents in dynamic adversarial games can significantly impact learning performance. However, the discussion of opponents in adversarial environments remains an area under exploration. In this paper, we propose a Step-level poliCy Optimization method through Play-And-Learn, SCO-PAL. Leveraging SCO-PAL, we conduct a detailed analysis of opponent selection by setting opponents at different levels and find that self-play is the most effective way to improve strategic reasoning in such adversarial environments. Utilizing SCO-PAL with self-play, we increase the average win rate against four opponents by approximately 30% compared to baselines and achieve a 54.76% win rate against GPT-4 in six adversarial games.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
Zhang, Yikai
Rong, Ye
Yuan, Siyu
Chen, Jiangjie
Xie, Jian
Xiao, Yanghua
Computation and Language
Existing language agents often encounter difficulties in dynamic adversarial games due to poor strategic reasoning. To mitigate this limitation, a promising approach is to allow agents to learn from game interactions automatically, without relying on costly expert-labeled data. Unlike static environments where agents receive fixed feedback or rewards, selecting appropriate opponents in dynamic adversarial games can significantly impact learning performance. However, the discussion of opponents in adversarial environments remains an area under exploration. In this paper, we propose a Step-level poliCy Optimization method through Play-And-Learn, SCO-PAL. Leveraging SCO-PAL, we conduct a detailed analysis of opponent selection by setting opponents at different levels and find that self-play is the most effective way to improve strategic reasoning in such adversarial environments. Utilizing SCO-PAL with self-play, we increase the average win rate against four opponents by approximately 30% compared to baselines and achieve a 54.76% win rate against GPT-4 in six adversarial games.
title Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
topic Computation and Language
url https://arxiv.org/abs/2510.16761