$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yaocheng, Zhu, Yuanheng, Chong, Wenyue, Tu, Songjun, Zhang, Qichao, Chai, Jiajun, Wang, Xiaohan, Lin, Wei, Yin, Guojun, Zhao, Dongbin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918520678252544
author Zhang, Yaocheng
Zhu, Yuanheng
Chong, Wenyue
Tu, Songjun
Zhang, Qichao
Chai, Jiajun
Wang, Xiaohan
Lin, Wei
Yin, Guojun
Zhao, Dongbin
author_facet Zhang, Yaocheng
Zhu, Yuanheng
Chong, Wenyue
Tu, Songjun
Zhang, Qichao
Chai, Jiajun
Wang, Xiaohan
Lin, Wei
Yin, Guojun
Zhao, Dongbin
contents Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In this work, we observe that self-play naturally produces a question construction path (QCP) during task generation, an intermediate artifact that captures the reverse solution process. This reveals a new source of privileged information: self-play can provide high-quality privileged information for the self-distillation at low cost and at scale, without relying on human feedback or curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play ($π$-Play), a novel multi-agent self-evolution framework combining self-play and self-distillation. In $π$-Play, an examiner generates tasks together with QCPs, and a teacher employs QCP as privileged context to densely supervise a student via self-distillation. This design transforms sparse-reward self-play into a dense-feedback co-evolution. Extensive experiments show that data-free $π$-Play surpasses fully supervised search agents and improves evolutionary efficiency by 2-3$\times$ over conventional self-play. Code is available at https://github.com/zhyaoch/pi-play.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14054
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle $π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
Zhang, Yaocheng
Zhu, Yuanheng
Chong, Wenyue
Tu, Songjun
Zhang, Qichao
Chai, Jiajun
Wang, Xiaohan
Lin, Wei
Yin, Guojun
Zhao, Dongbin
Machine Learning
Computation and Language
Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In this work, we observe that self-play naturally produces a question construction path (QCP) during task generation, an intermediate artifact that captures the reverse solution process. This reveals a new source of privileged information: self-play can provide high-quality privileged information for the self-distillation at low cost and at scale, without relying on human feedback or curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play ($π$-Play), a novel multi-agent self-evolution framework combining self-play and self-distillation. In $π$-Play, an examiner generates tasks together with QCPs, and a teacher employs QCP as privileged context to densely supervise a student via self-distillation. This design transforms sparse-reward self-play into a dense-feedback co-evolution. Extensive experiments show that data-free $π$-Play surpasses fully supervised search agents and improves evolutionary efficiency by 2-3$\times$ over conventional self-play. Code is available at https://github.com/zhyaoch/pi-play.
title $π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2604.14054