Rate optimal learning of equilibria from data
Fuente:
arXiv
Saved in:
| Main Authors: | Freihaut, Till, Viano, Luca, Nevali, Emanuele, Cevher, Volkan, Geist, Matthieu, Ramponi, Giorgia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
by: Viano, Luca, et al.
Published: (2026)
by: Viano, Luca, et al.
Published: (2026)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
by: Freihaut, Till, et al.
Published: (2025)
by: Freihaut, Till, et al.
Published: (2025)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
by: Barla, Adam, et al.
Published: (2026)
by: Barla, Adam, et al.
Published: (2026)
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
by: Freihaut, Till, et al.
Published: (2024)
by: Freihaut, Till, et al.
Published: (2024)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)
by: Viel, Stefano, et al.
Published: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Clustered KL-barycenter design for policy evaluation
by: Weissmann, Simon, et al.
Published: (2025)
by: Weissmann, Simon, et al.
Published: (2025)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
by: Macuglia, Maël, et al.
Published: (2025)
by: Macuglia, Maël, et al.
Published: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
Best of Both Worlds: Regret Minimization versus Minimax Play
by: Müller, Adrian, et al.
Published: (2025)
by: Müller, Adrian, et al.
Published: (2025)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
by: Viano, Luca, et al.
Published: (2026)
by: Viano, Luca, et al.
Published: (2026)
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)
by: Afzal, Arshia, et al.
Published: (2024)
Addressing Label Shift in Distributed Learning via Entropy Regularization
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
by: Baur, Raphaël, et al.
Published: (2026)
by: Baur, Raphaël, et al.
Published: (2026)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
by: Thoma, Vinzenz, et al.
Published: (2024)
by: Thoma, Vinzenz, et al.
Published: (2024)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
by: Koren, Uri, et al.
Published: (2025)
by: Koren, Uri, et al.
Published: (2025)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Bootstrapping Expectiles in Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2024)
by: Clavier, Pierre, et al.
Published: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
by: Wu, Yongtao, et al.
Published: (2024)
by: Wu, Yongtao, et al.
Published: (2024)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2025)
by: Kumar, Navdeep, et al.
Published: (2025)
Aligning Language Models from User Interactions
by: Buening, Thomas Kleine, et al.
Published: (2026)
by: Buening, Thomas Kleine, et al.
Published: (2026)
Space Robotics Bench: Robot Learning Beyond Earth
by: Orsula, Andrej, et al.
Published: (2025)
by: Orsula, Andrej, et al.
Published: (2025)
Learning Tool-Aware Adaptive Compliant Control for Autonomous Regolith Excavation
by: Orsula, Andrej, et al.
Published: (2025)
by: Orsula, Andrej, et al.
Published: (2025)
Sim2Dust: Mastering Dynamic Waypoint Tracking on Granular Media
by: Orsula, Andrej, et al.
Published: (2025)
by: Orsula, Andrej, et al.
Published: (2025)
Leveraging Procedural Generation for Learning Autonomous Peg-in-Hole Assembly in Space
by: Orsula, Andrej, et al.
Published: (2024)
by: Orsula, Andrej, et al.
Published: (2024)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
by: Candogan, Leyla Naz, et al.
Published: (2025)
by: Candogan, Leyla Naz, et al.
Published: (2025)
Efficient local linearity regularization to overcome catastrophic overfitting
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Similar Items
-
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
by: Viano, Luca, et al.
Published: (2026) -
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
by: Freihaut, Till, et al.
Published: (2025) -
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
by: Barla, Adam, et al.
Published: (2026) -
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
by: Freihaut, Till, et al.
Published: (2024) -
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)