Saved in:
| Main Authors: | Li, Shangzhe, Zhang, Xuchao, Bansal, Chetan, Zhang, Weitong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.01357 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Imitation from Observations with Trajectory-Level Generative Embeddings
by: Qu, Yongtao, et al.
Published: (2026)
by: Qu, Yongtao, et al.
Published: (2026)
CREAM: Consistency Regularized Self-Rewarding Language Models
by: Wang, Zhaoyang, et al.
Published: (2024)
by: Wang, Zhaoyang, et al.
Published: (2024)
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026)
by: Yu, Tianrun, et al.
Published: (2026)
Reward-free World Models for Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2024)
by: Li, Shangzhe, et al.
Published: (2024)
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Auto-Encoding Adversarial Imitation Learning
by: Zhang, Kaifeng, et al.
Published: (2022)
by: Zhang, Kaifeng, et al.
Published: (2022)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Multi-Agent Generative Adversarial Interactive Self-Imitation Learning for AUV Formation Control and Obstacle Avoidance
by: Fang, Zheng, et al.
Published: (2024)
by: Fang, Zheng, et al.
Published: (2024)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
XQSV: A Structurally Variable Network to Imitate Human Play in Xiangqi
by: Zhou, Chenliang
Published: (2024)
by: Zhou, Chenliang
Published: (2024)
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control
by: Ye, Yutong, et al.
Published: (2025)
by: Ye, Yutong, et al.
Published: (2025)
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
by: Xu, Tian, et al.
Published: (2026)
by: Xu, Tian, et al.
Published: (2026)
R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Latent Wasserstein Adversarial Imitation Learning
by: Yang, Siqi, et al.
Published: (2026)
by: Yang, Siqi, et al.
Published: (2026)
Agnostic Interactive Imitation Learning: New Theory and Practical Algorithms
by: Li, Yichen, et al.
Published: (2023)
by: Li, Yichen, et al.
Published: (2023)
A Theoretical Framework for Self-Play Theorem Proving Algorithms
by: Chen, Thomas, et al.
Published: (2026)
by: Chen, Thomas, et al.
Published: (2026)
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Provable Memory Efficient Self-Play Algorithm for Model-free Reinforcement Learning
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
Scaling Self-Play with Self-Guidance
by: Bailey, Luke, et al.
Published: (2026)
by: Bailey, Luke, et al.
Published: (2026)
Self-evolved Imitation Learning in Simulated World
by: Ye, Yifan, et al.
Published: (2025)
by: Ye, Yifan, et al.
Published: (2025)
AutoAdapt: An Automated Domain Adaptation Framework for LLMs
by: Sinha, Sidharth, et al.
Published: (2026)
by: Sinha, Sidharth, et al.
Published: (2026)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
Adversarial Imitation Learning via Boosting
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Sample-efficient Adversarial Imitation Learning
by: Jung, Dahuin, et al.
Published: (2023)
by: Jung, Dahuin, et al.
Published: (2023)
Learning to Drive via Asymmetric Self-Play
by: Zhang, Chris, et al.
Published: (2024)
by: Zhang, Chris, et al.
Published: (2024)
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
by: Xu, Tian, et al.
Published: (2022)
by: Xu, Tian, et al.
Published: (2022)
Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
by: Zhuang, Zifeng, et al.
Published: (2025)
by: Zhuang, Zifeng, et al.
Published: (2025)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
Exploring LLM-based Agents for Root Cause Analysis
by: Roy, Devjeet, et al.
Published: (2024)
by: Roy, Devjeet, et al.
Published: (2024)
MILES: Making Imitation Learning Easy with Self-Supervision
by: Papagiannis, Georgios, et al.
Published: (2024)
by: Papagiannis, Georgios, et al.
Published: (2024)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
by: Xu, Tian, et al.
Published: (2024)
by: Xu, Tian, et al.
Published: (2024)
Dexterous Manipulation through Imitation Learning: A Survey
by: An, Shan, et al.
Published: (2025)
by: An, Shan, et al.
Published: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Similar Items
-
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025) -
Imitation from Observations with Trajectory-Level Generative Embeddings
by: Qu, Yongtao, et al.
Published: (2026) -
CREAM: Consistency Regularized Self-Rewarding Language Models
by: Wang, Zhaoyang, et al.
Published: (2024) -
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026) -
Reward-free World Models for Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2024)