Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
Fuente:
arXiv
Guardado en:
| Autores principales: | Sheebaelhamd, Ziyad, Viano, Luca, Cevher, Volkan, Vernade, Claire |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)
por: Viano, Luca, et al.
Publicado: (2024)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
por: Viel, Stefano, et al.
Publicado: (2025)
por: Viel, Stefano, et al.
Publicado: (2025)
Quantization-Free Autoregressive Action Transformer
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2025)
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2025)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
por: Barla, Adam, et al.
Publicado: (2026)
por: Barla, Adam, et al.
Publicado: (2026)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
por: Moulin, Antoine, et al.
Publicado: (2025)
por: Moulin, Antoine, et al.
Publicado: (2025)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
Best of Both Worlds: Regret Minimization versus Minimax Play
por: Müller, Adrian, et al.
Publicado: (2025)
por: Müller, Adrian, et al.
Publicado: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
por: Viano, Luca, et al.
Publicado: (2026)
por: Viano, Luca, et al.
Publicado: (2026)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
por: Bergerault, Antoine, et al.
Publicado: (2026)
por: Bergerault, Antoine, et al.
Publicado: (2026)
Rate optimal learning of equilibria from data
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
por: Wu, Yongtao, et al.
Publicado: (2025)
por: Wu, Yongtao, et al.
Publicado: (2025)
Efficient Personalization of Generative Models via Optimal Experimental Design
por: Schacht, Guy, et al.
Publicado: (2025)
por: Schacht, Guy, et al.
Publicado: (2025)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
por: Moulin, Antoine, et al.
Publicado: (2025)
por: Moulin, Antoine, et al.
Publicado: (2025)
Tight Sample Complexity Bounds for Entropic Best Policy Identification
por: Essakine, Amer, et al.
Publicado: (2026)
por: Essakine, Amer, et al.
Publicado: (2026)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
por: Xie, Wanyun, et al.
Publicado: (2026)
por: Xie, Wanyun, et al.
Publicado: (2026)
Efficient Large Language Model Inference with Neural Block Linearization
por: Erdogan, Mete, et al.
Publicado: (2025)
por: Erdogan, Mete, et al.
Publicado: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
por: Afzal, Arshia, et al.
Publicado: (2025)
por: Afzal, Arshia, et al.
Publicado: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
por: Liu, Fanghui, et al.
Publicado: (2024)
por: Liu, Fanghui, et al.
Publicado: (2024)
SAMPa: Sharpness-aware Minimization Parallelized
por: Xie, Wanyun, et al.
Publicado: (2024)
por: Xie, Wanyun, et al.
Publicado: (2024)
Commit to the Bit: Reactive Reinforcement Learning Done Right
por: Eberhard, Onno, et al.
Publicado: (2026)
por: Eberhard, Onno, et al.
Publicado: (2026)
Partially Observable Reinforcement Learning with Memory Traces
por: Eberhard, Onno, et al.
Publicado: (2025)
por: Eberhard, Onno, et al.
Publicado: (2025)
Non-Stationary Lipschitz Bandits
por: Nguyen, Nicolas, et al.
Publicado: (2025)
por: Nguyen, Nicolas, et al.
Publicado: (2025)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
por: Oikonomidis, Konstantinos, et al.
Publicado: (2026)
por: Oikonomidis, Konstantinos, et al.
Publicado: (2026)
Variational Bayes Portfolio Construction
por: Nguyen, Nicolas, et al.
Publicado: (2024)
por: Nguyen, Nicolas, et al.
Publicado: (2024)
Efficient Continual Finite-Sum Minimization
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
por: Xie, Wanyun, et al.
Publicado: (2025)
por: Xie, Wanyun, et al.
Publicado: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
por: Pethick, Thomas, et al.
Publicado: (2023)
por: Pethick, Thomas, et al.
Publicado: (2023)
A Pontryagin Perspective on Reinforcement Learning
por: Eberhard, Onno, et al.
Publicado: (2024)
por: Eberhard, Onno, et al.
Publicado: (2024)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
por: Afzal, Arshia, et al.
Publicado: (2024)
por: Afzal, Arshia, et al.
Publicado: (2024)
Efficient Risk-sensitive Planning via Entropic Risk Measures
por: Marthe, Alexandre, et al.
Publicado: (2025)
por: Marthe, Alexandre, et al.
Publicado: (2025)
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
por: Cao, Linfeng, et al.
Publicado: (2025)
por: Cao, Linfeng, et al.
Publicado: (2025)
Optimistic Dual Averaging Unifies Modern Optimizers
por: Pethick, Thomas, et al.
Publicado: (2026)
por: Pethick, Thomas, et al.
Publicado: (2026)
Easy Data Unlearning Bench
por: Rinberg, Roy, et al.
Publicado: (2026)
por: Rinberg, Roy, et al.
Publicado: (2026)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
por: Chen, Yihang, et al.
Publicado: (2024)
por: Chen, Yihang, et al.
Publicado: (2024)
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
por: Nguyen, Nicolas, et al.
Publicado: (2024)
por: Nguyen, Nicolas, et al.
Publicado: (2024)
Online Decision Deferral under Budget Constraints
por: Reid, Mirabel, et al.
Publicado: (2024)
por: Reid, Mirabel, et al.
Publicado: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
por: Bal, Melis Ilayda, et al.
Publicado: (2025)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
por: Balef, Amir Rezaei, et al.
Publicado: (2025)
por: Balef, Amir Rezaei, et al.
Publicado: (2025)
Ejemplares similares
-
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025) -
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024) -
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
por: Viel, Stefano, et al.
Publicado: (2025) -
Quantization-Free Autoregressive Action Transformer
por: Sheebaelhamd, Ziyad, et al.
Publicado: (2025) -
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
por: Barla, Adam, et al.
Publicado: (2026)