Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Minzheng, Luo, Run, Wang, Yanbo, Liu, Zichen, Tan, Yuqiao, Tan, Tao, Nan, Xu, Zheng, Yinhe, Mao, Wenji
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918492382429184
author Wang, Minzheng
Luo, Run
Wang, Yanbo
Liu, Zichen
Tan, Yuqiao
Tan, Tao
Nan, Xu
Zheng, Yinhe
Mao, Wenji
author_facet Wang, Minzheng
Luo, Run
Wang, Yanbo
Liu, Zichen
Tan, Yuqiao
Tan, Tao
Nan, Xu
Zheng, Yinhe
Mao, Wenji
contents While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08721
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
Wang, Minzheng
Luo, Run
Wang, Yanbo
Liu, Zichen
Tan, Yuqiao
Tan, Tao
Nan, Xu
Zheng, Yinhe
Mao, Wenji
Computation and Language
While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.
title Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
topic Computation and Language
url https://arxiv.org/abs/2605.08721