$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918520678252544 |
|---|---|
| author | Zhang, Yaocheng Zhu, Yuanheng Chong, Wenyue Tu, Songjun Zhang, Qichao Chai, Jiajun Wang, Xiaohan Lin, Wei Yin, Guojun Zhao, Dongbin |
| author_facet | Zhang, Yaocheng Zhu, Yuanheng Chong, Wenyue Tu, Songjun Zhang, Qichao Chai, Jiajun Wang, Xiaohan Lin, Wei Yin, Guojun Zhao, Dongbin |
| contents | Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In this work, we observe that self-play naturally produces a question construction path (QCP) during task generation, an intermediate artifact that captures the reverse solution process. This reveals a new source of privileged information: self-play can provide high-quality privileged information for the self-distillation at low cost and at scale, without relying on human feedback or curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play ($π$-Play), a novel multi-agent self-evolution framework combining self-play and self-distillation. In $π$-Play, an examiner generates tasks together with QCPs, and a teacher employs QCP as privileged context to densely supervise a student via self-distillation. This design transforms sparse-reward self-play into a dense-feedback co-evolution. Extensive experiments show that data-free $π$-Play surpasses fully supervised search agents and improves evolutionary efficiency by 2-3$\times$ over conventional self-play. Code is available at https://github.com/zhyaoch/pi-play. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_14054 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | $π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data Zhang, Yaocheng Zhu, Yuanheng Chong, Wenyue Tu, Songjun Zhang, Qichao Chai, Jiajun Wang, Xiaohan Lin, Wei Yin, Guojun Zhao, Dongbin Machine Learning Computation and Language Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due to sparse rewards, weak credit assignment, and limited labeled data. Self-play offers a scalable route to reduce data dependence, but conventional self-play optimizes students only through sparse outcome rewards, leading to low learning efficiency. In this work, we observe that self-play naturally produces a question construction path (QCP) during task generation, an intermediate artifact that captures the reverse solution process. This reveals a new source of privileged information: self-play can provide high-quality privileged information for the self-distillation at low cost and at scale, without relying on human feedback or curated privileged information. Leveraging this insight, we propose Privileged Information Self-Play ($π$-Play), a novel multi-agent self-evolution framework combining self-play and self-distillation. In $π$-Play, an examiner generates tasks together with QCPs, and a teacher employs QCP as privileged context to densely supervise a student via self-distillation. This design transforms sparse-reward self-play into a dense-feedback co-evolution. Extensive experiments show that data-free $π$-Play surpasses fully supervised search agents and improves evolutionary efficiency by 2-3$\times$ over conventional self-play. Code is available at https://github.com/zhyaoch/pi-play. |
| title | $π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2604.14054 |