Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Shangzhe, Zhang, Xuchao, Bansal, Chetan, Zhang, Weitong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
Imitation from Observations with Trajectory-Level Generative Embeddings
di: Qu, Yongtao, et al.
Pubblicazione: (2026)
di: Qu, Yongtao, et al.
Pubblicazione: (2026)
CREAM: Consistency Regularized Self-Rewarding Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024)
Reward-free World Models for Online Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2024)
di: Li, Shangzhe, et al.
Pubblicazione: (2024)
Learning Robust Reasoning through Guided Adversarial Self-Play
di: Li, Shuozhe, et al.
Pubblicazione: (2026)
di: Li, Shuozhe, et al.
Pubblicazione: (2026)
Provable and Practical In-Context Policy Optimization for Self-Improvement
di: Yu, Tianrun, et al.
Pubblicazione: (2026)
di: Yu, Tianrun, et al.
Pubblicazione: (2026)
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
di: Li, Shangzhe, et al.
Pubblicazione: (2026)
di: Li, Shangzhe, et al.
Pubblicazione: (2026)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
di: Gupta, Taneesh, et al.
Pubblicazione: (2025)
di: Gupta, Taneesh, et al.
Pubblicazione: (2025)
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
Auto-Encoding Adversarial Imitation Learning
di: Zhang, Kaifeng, et al.
Pubblicazione: (2022)
di: Zhang, Kaifeng, et al.
Pubblicazione: (2022)
Multi-Agent Generative Adversarial Interactive Self-Imitation Learning for AUV Formation Control and Obstacle Avoidance
di: Fang, Zheng, et al.
Pubblicazione: (2024)
di: Fang, Zheng, et al.
Pubblicazione: (2024)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
di: Tan, Zhewen, et al.
Pubblicazione: (2026)
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
di: Xu, Tian, et al.
Pubblicazione: (2026)
di: Xu, Tian, et al.
Pubblicazione: (2026)
XQSV: A Structurally Variable Network to Imitate Human Play in Xiangqi
di: Zhou, Chenliang
Pubblicazione: (2024)
di: Zhou, Chenliang
Pubblicazione: (2024)
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control
di: Ye, Yutong, et al.
Pubblicazione: (2025)
di: Ye, Yutong, et al.
Pubblicazione: (2025)
Agnostic Interactive Imitation Learning: New Theory and Practical Algorithms
di: Li, Yichen, et al.
Pubblicazione: (2023)
di: Li, Yichen, et al.
Pubblicazione: (2023)
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
di: Li, Gengsheng, et al.
Pubblicazione: (2026)
Latent Wasserstein Adversarial Imitation Learning
di: Yang, Siqi, et al.
Pubblicazione: (2026)
di: Yang, Siqi, et al.
Pubblicazione: (2026)
A Theoretical Framework for Self-Play Theorem Proving Algorithms
di: Chen, Thomas, et al.
Pubblicazione: (2026)
di: Chen, Thomas, et al.
Pubblicazione: (2026)
Self-Improving AI Agents through Self-Play
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
di: Chojecki, Przemyslaw
Pubblicazione: (2025)
Provable Memory Efficient Self-Play Algorithm for Model-free Reinforcement Learning
di: Li, Na, et al.
Pubblicazione: (2025)
di: Li, Na, et al.
Pubblicazione: (2025)
Scaling Self-Play with Self-Guidance
di: Bailey, Luke, et al.
Pubblicazione: (2026)
di: Bailey, Luke, et al.
Pubblicazione: (2026)
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
di: Xu, Tian, et al.
Pubblicazione: (2022)
di: Xu, Tian, et al.
Pubblicazione: (2022)
Self-evolved Imitation Learning in Simulated World
di: Ye, Yifan, et al.
Pubblicazione: (2025)
di: Ye, Yifan, et al.
Pubblicazione: (2025)
Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning
di: Li, Na, et al.
Pubblicazione: (2025)
di: Li, Na, et al.
Pubblicazione: (2025)
Adversarial Imitation Learning via Boosting
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
Sample-efficient Adversarial Imitation Learning
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
di: Zhang, Yaocheng, et al.
Pubblicazione: (2026)
di: Zhang, Yaocheng, et al.
Pubblicazione: (2026)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
di: Zhuang, Zifeng, et al.
Pubblicazione: (2025)
di: Zhuang, Zifeng, et al.
Pubblicazione: (2025)
Learning to Drive via Asymmetric Self-Play
di: Zhang, Chris, et al.
Pubblicazione: (2024)
di: Zhang, Chris, et al.
Pubblicazione: (2024)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
di: Wang, Lu, et al.
Pubblicazione: (2024)
di: Wang, Lu, et al.
Pubblicazione: (2024)
Diffusion-Reward Adversarial Imitation Learning
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
di: Xu, Tian, et al.
Pubblicazione: (2024)
di: Xu, Tian, et al.
Pubblicazione: (2024)
MILES: Making Imitation Learning Easy with Self-Supervision
di: Papagiannis, Georgios, et al.
Pubblicazione: (2024)
di: Papagiannis, Georgios, et al.
Pubblicazione: (2024)
Dexterous Manipulation through Imitation Learning: A Survey
di: An, Shan, et al.
Pubblicazione: (2025)
di: An, Shan, et al.
Pubblicazione: (2025)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
di: Chaudhary, Gaurav, et al.
Pubblicazione: (2025)
di: Chaudhary, Gaurav, et al.
Pubblicazione: (2025)
How to Train Your Robots? The Impact of Demonstration Modality on Imitation Learning
di: Li, Haozhuo, et al.
Pubblicazione: (2025)
di: Li, Haozhuo, et al.
Pubblicazione: (2025)
Match or Replay: Self Imitating Proximal Policy Optimization
di: Chaudhary, Gaurav, et al.
Pubblicazione: (2026)
di: Chaudhary, Gaurav, et al.
Pubblicazione: (2026)
REFA: Reference Free Alignment for multi-preference optimization
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2025) -
Imitation from Observations with Trajectory-Level Generative Embeddings
di: Qu, Yongtao, et al.
Pubblicazione: (2026) -
CREAM: Consistency Regularized Self-Rewarding Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024) -
Reward-free World Models for Online Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2024) -
Learning Robust Reasoning through Guided Adversarial Self-Play
di: Li, Shuozhe, et al.
Pubblicazione: (2026)