Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Tian, Wang, Chenyang, Zhai, Xiaochen, Li, Ziniu, Li, Yi-Chen, Yu, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
di: Xu, Tian, et al.
Pubblicazione: (2022)
di: Xu, Tian, et al.
Pubblicazione: (2022)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
di: Xu, Tian, et al.
Pubblicazione: (2024)
di: Xu, Tian, et al.
Pubblicazione: (2024)
Bellman Error Centering
di: Chen, Xingguo, et al.
Pubblicazione: (2025)
di: Chen, Xingguo, et al.
Pubblicazione: (2025)
Provably Efficient Off-Policy Adversarial Imitation Learning with Convergence Guarantees
di: Chen, Yilei, et al.
Pubblicazione: (2024)
di: Chen, Yilei, et al.
Pubblicazione: (2024)
Policy Optimization in RLHF: The Impact of Out-of-preference Data
di: Li, Ziniu, et al.
Pubblicazione: (2023)
di: Li, Ziniu, et al.
Pubblicazione: (2023)
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
di: Xu, Tian, et al.
Pubblicazione: (2026)
di: Xu, Tian, et al.
Pubblicazione: (2026)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
di: Li, Yingru, et al.
Pubblicazione: (2025)
di: Li, Yingru, et al.
Pubblicazione: (2025)
Latent Wasserstein Adversarial Imitation Learning
di: Yang, Siqi, et al.
Pubblicazione: (2026)
di: Yang, Siqi, et al.
Pubblicazione: (2026)
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
di: Li, Yichen, et al.
Pubblicazione: (2024)
di: Li, Yichen, et al.
Pubblicazione: (2024)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
di: Golowich, Noah, et al.
Pubblicazione: (2024)
di: Golowich, Noah, et al.
Pubblicazione: (2024)
One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL
di: Chen, Elynn, et al.
Pubblicazione: (2026)
di: Chen, Elynn, et al.
Pubblicazione: (2026)
Auto-Encoding Adversarial Imitation Learning
di: Zhang, Kaifeng, et al.
Pubblicazione: (2022)
di: Zhang, Kaifeng, et al.
Pubblicazione: (2022)
Diffusion-Reward Adversarial Imitation Learning
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
di: Lai, Chun-Mao, et al.
Pubblicazione: (2024)
Towards Optimal Adversarial Robust Q-learning with Bellman Infinity-error
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
Off-Policy Value-Based Reinforcement Learning for Large Language Models
di: Wang, Peng-Yuan, et al.
Pubblicazione: (2026)
di: Wang, Peng-Yuan, et al.
Pubblicazione: (2026)
Provable Unrestricted Adversarial Training without Compromise with Generalizability
di: Zhang, Lilin, et al.
Pubblicazione: (2023)
di: Zhang, Lilin, et al.
Pubblicazione: (2023)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
di: Freihaut, Till, et al.
Pubblicazione: (2025)
di: Freihaut, Till, et al.
Pubblicazione: (2025)
Deterministic Exploration via Stationary Bellman Error Maximization
di: Griesbach, Sebastian, et al.
Pubblicazione: (2024)
di: Griesbach, Sebastian, et al.
Pubblicazione: (2024)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
di: Cho, Taehyun, et al.
Pubblicazione: (2024)
di: Cho, Taehyun, et al.
Pubblicazione: (2024)
Looped Transformers with Layer Normalization Provably Learn the Power Method
di: Wu, Lyumin, et al.
Pubblicazione: (2026)
di: Wu, Lyumin, et al.
Pubblicazione: (2026)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2026)
di: Li, Shangzhe, et al.
Pubblicazione: (2026)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
di: Moulin, Antoine, et al.
Pubblicazione: (2025)
di: Moulin, Antoine, et al.
Pubblicazione: (2025)
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
di: Li, Shangzhe, et al.
Pubblicazione: (2025)
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
di: Li, Ziniu, et al.
Pubblicazione: (2023)
di: Li, Ziniu, et al.
Pubblicazione: (2023)
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
di: Jia, Chengxing, et al.
Pubblicazione: (2024)
di: Jia, Chengxing, et al.
Pubblicazione: (2024)
PAIL: Performance based Adversarial Imitation Learning Engine for Carbon Neutral Optimization
di: Ye, Yuyang, et al.
Pubblicazione: (2024)
di: Ye, Yuyang, et al.
Pubblicazione: (2024)
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
di: Li, Binghui, et al.
Pubblicazione: (2024)
di: Li, Binghui, et al.
Pubblicazione: (2024)
SD2AIL: Adversarial Imitation Learning from Synthetic Demonstrations via Diffusion Models
di: Li, Pengcheng, et al.
Pubblicazione: (2025)
di: Li, Pengcheng, et al.
Pubblicazione: (2025)
Adversarial Imitation Learning via Boosting
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
Sample-efficient Adversarial Imitation Learning
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
di: Jung, Dahuin, et al.
Pubblicazione: (2023)
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
di: Omura, Motoki, et al.
Pubblicazione: (2024)
di: Omura, Motoki, et al.
Pubblicazione: (2024)
Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach
di: Lu, Chenbei, et al.
Pubblicazione: (2025)
di: Lu, Chenbei, et al.
Pubblicazione: (2025)
MICRO: Model-Based Offline Reinforcement Learning with a Conservative Bellman Operator
di: Liu, Xiao-Yin, et al.
Pubblicazione: (2023)
di: Liu, Xiao-Yin, et al.
Pubblicazione: (2023)
Unlocking TriLevel Learning with Level-Wise Zeroth Order Constraints: Distributed Algorithms and Provable Non-Asymptotic Convergence
di: Jiao, Yang, et al.
Pubblicazione: (2024)
di: Jiao, Yang, et al.
Pubblicazione: (2024)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
di: Yang, Chen, et al.
Pubblicazione: (2024)
di: Yang, Chen, et al.
Pubblicazione: (2024)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
di: Zhong, Han, et al.
Pubblicazione: (2021)
di: Zhong, Han, et al.
Pubblicazione: (2021)
PAGAR: Taming Reward Misalignment in Inverse Reinforcement Learning-Based Imitation Learning with Protagonist Antagonist Guided Adversarial Reward
di: Zhou, Weichao, et al.
Pubblicazione: (2023)
di: Zhou, Weichao, et al.
Pubblicazione: (2023)
Ranking-based Client Selection with Imitation Learning for Efficient Federated Learning
di: Tian, Chunlin, et al.
Pubblicazione: (2024)
di: Tian, Chunlin, et al.
Pubblicazione: (2024)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
di: Patterson, Andrew, et al.
Pubblicazione: (2021)
di: Patterson, Andrew, et al.
Pubblicazione: (2021)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
di: Sheebaelhamd, Ziyad, et al.
Pubblicazione: (2026)
di: Sheebaelhamd, Ziyad, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
di: Xu, Tian, et al.
Pubblicazione: (2022) -
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
di: Xu, Tian, et al.
Pubblicazione: (2024) -
Bellman Error Centering
di: Chen, Xingguo, et al.
Pubblicazione: (2025) -
Provably Efficient Off-Policy Adversarial Imitation Learning with Convergence Guarantees
di: Chen, Yilei, et al.
Pubblicazione: (2024) -
Policy Optimization in RLHF: The Impact of Out-of-preference Data
di: Li, Ziniu, et al.
Pubblicazione: (2023)