Multi-agent imitation learning with function approximation: Linear Markov games and beyond
Fuente:
arXiv
Guardado en:
| Autores principales: | Viano, Luca, Freihaut, Till, Nevali, Emanuele, Cevher, Volkan, Geist, Matthieu, Ramponi, Giorgia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
por: Barla, Adam, et al.
Publicado: (2026)
por: Barla, Adam, et al.
Publicado: (2026)
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
por: Freihaut, Till, et al.
Publicado: (2024)
por: Freihaut, Till, et al.
Publicado: (2024)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)
por: Viano, Luca, et al.
Publicado: (2024)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
por: Viel, Stefano, et al.
Publicado: (2025)
por: Viel, Stefano, et al.
Publicado: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2026)
Clustered KL-barycenter design for policy evaluation
por: Weissmann, Simon, et al.
Publicado: (2025)
por: Weissmann, Simon, et al.
Publicado: (2025)
Truly No-Regret Learning in Constrained MDPs
por: Müller, Adrian, et al.
Publicado: (2024)
por: Müller, Adrian, et al.
Publicado: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
por: Wu, Yongtao, et al.
Publicado: (2025)
por: Wu, Yongtao, et al.
Publicado: (2025)
Best of Both Worlds: Regret Minimization versus Minimax Play
por: Müller, Adrian, et al.
Publicado: (2025)
por: Müller, Adrian, et al.
Publicado: (2025)
Periodic agent-state based Q-learning for POMDPs
por: Sinha, Amit, et al.
Publicado: (2024)
por: Sinha, Amit, et al.
Publicado: (2024)
Convergence of regularized agent-state-based Q-learning in POMDPs
por: Sinha, Amit, et al.
Publicado: (2025)
por: Sinha, Amit, et al.
Publicado: (2025)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
por: Pethick, Thomas, et al.
Publicado: (2023)
por: Pethick, Thomas, et al.
Publicado: (2023)
Efficient Large Language Model Inference with Neural Block Linearization
por: Erdogan, Mete, et al.
Publicado: (2025)
por: Erdogan, Mete, et al.
Publicado: (2025)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
por: Xie, Wanyun, et al.
Publicado: (2026)
por: Xie, Wanyun, et al.
Publicado: (2026)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
por: Liu, Fanghui, et al.
Publicado: (2024)
por: Liu, Fanghui, et al.
Publicado: (2024)
SAMPa: Sharpness-aware Minimization Parallelized
por: Xie, Wanyun, et al.
Publicado: (2024)
por: Xie, Wanyun, et al.
Publicado: (2024)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
por: Bergerault, Antoine, et al.
Publicado: (2026)
por: Bergerault, Antoine, et al.
Publicado: (2026)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
por: Macuglia, Maël, et al.
Publicado: (2025)
por: Macuglia, Maël, et al.
Publicado: (2025)
Learning to Remove Cuts in Integer Linear Programming
por: Puigdemont, Pol, et al.
Publicado: (2024)
por: Puigdemont, Pol, et al.
Publicado: (2024)
Hadamard product in deep learning: Introduction, Advances and Challenges
por: Chrysos, Grigorios G, et al.
Publicado: (2025)
por: Chrysos, Grigorios G, et al.
Publicado: (2025)
Robust NAS under adversarial training: benchmark, theory, and beyond
por: Wu, Yongtao, et al.
Publicado: (2024)
por: Wu, Yongtao, et al.
Publicado: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
por: Xie, Wanyun, et al.
Publicado: (2025)
por: Xie, Wanyun, et al.
Publicado: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
por: Kumar, Navdeep, et al.
Publicado: (2025)
por: Kumar, Navdeep, et al.
Publicado: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
por: Afzal, Arshia, et al.
Publicado: (2025)
por: Afzal, Arshia, et al.
Publicado: (2025)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
por: Moulin, Antoine, et al.
Publicado: (2025)
por: Moulin, Antoine, et al.
Publicado: (2025)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
por: Moulin, Antoine, et al.
Publicado: (2025)
por: Moulin, Antoine, et al.
Publicado: (2025)
Optimistic Dual Averaging Unifies Modern Optimizers
por: Pethick, Thomas, et al.
Publicado: (2026)
por: Pethick, Thomas, et al.
Publicado: (2026)
Easy Data Unlearning Bench
por: Rinberg, Roy, et al.
Publicado: (2026)
por: Rinberg, Roy, et al.
Publicado: (2026)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
por: Chen, Yihang, et al.
Publicado: (2024)
por: Chen, Yihang, et al.
Publicado: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
por: Kamoutsi, Angeliki, et al.
Publicado: (2024)
por: Kamoutsi, Angeliki, et al.
Publicado: (2024)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
por: Cheng, Yixin, et al.
Publicado: (2024)
por: Cheng, Yixin, et al.
Publicado: (2024)
Preference Elicitation for Offline Reinforcement Learning
por: Pace, Alizée, et al.
Publicado: (2024)
por: Pace, Alizée, et al.
Publicado: (2024)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
por: Haas, Moritz, et al.
Publicado: (2024)
por: Haas, Moritz, et al.
Publicado: (2024)
Solving robust MDPs as a sequence of static RL problems
por: Zouitine, Adil, et al.
Publicado: (2024)
por: Zouitine, Adil, et al.
Publicado: (2024)
Accelerating Spectral Clustering under Fairness Constraints
por: Tonin, Francesco, et al.
Publicado: (2025)
por: Tonin, Francesco, et al.
Publicado: (2025)
Ejemplares similares
-
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025) -
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025) -
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
por: Barla, Adam, et al.
Publicado: (2026) -
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
por: Freihaut, Till, et al.
Publicado: (2024) -
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)