Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yongtao, Viano, Luca, Chen, Yihang, Zhu, Zhenyu, Antonakopoulos, Kimon, Gu, Quanquan, Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Universal Gradient Methods for Stochastic Convex Optimization
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
On the Generalization of Stochastic Gradient Descent with Momentum
von: Ramezani-Kebrya, Ali, et al.
Veröffentlicht: (2018)
von: Ramezani-Kebrya, Ali, et al.
Veröffentlicht: (2018)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
Polynomial Convergence of Bandit No-Regret Dynamics in Congestion Games
von: Dadi, Leello, et al.
Veröffentlicht: (2024)
von: Dadi, Leello, et al.
Veröffentlicht: (2024)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
Training Deep Learning Models with Norm-Constrained LMOs
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
von: Viano, Luca, et al.
Veröffentlicht: (2024)
von: Viano, Luca, et al.
Veröffentlicht: (2024)
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
Training Neural Networks at Any Scale
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
von: Barla, Adam, et al.
Veröffentlicht: (2026)
von: Barla, Adam, et al.
Veröffentlicht: (2026)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
Advancing the lower bounds: An accelerated, stochastic, second-order method with optimal adaptation to inexactness
von: Agafonov, Artem, et al.
Veröffentlicht: (2023)
von: Agafonov, Artem, et al.
Veröffentlicht: (2023)
Best of Both Worlds: Regret Minimization versus Minimax Play
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
Optimistic Dual Averaging Unifies Modern Optimizers
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
Rate optimal learning of equilibria from data
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
Membership Inference Attacks against Large Vision-Language Models
von: Li, Zhan, et al.
Veröffentlicht: (2024)
von: Li, Zhan, et al.
Veröffentlicht: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
Convergence of Markov Chains for Constant Step-size Stochastic Gradient Descent with Separable Functions
von: Shirokoff, David, et al.
Veröffentlicht: (2024)
von: Shirokoff, David, et al.
Veröffentlicht: (2024)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
von: Fu, Jia, et al.
Veröffentlicht: (2025)
von: Fu, Jia, et al.
Veröffentlicht: (2025)
Hadamard product in deep learning: Introduction, Advances and Challenges
von: Chrysos, Grigorios G, et al.
Veröffentlicht: (2025)
von: Chrysos, Grigorios G, et al.
Veröffentlicht: (2025)
Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games
von: Dong, Jing, et al.
Veröffentlicht: (2024)
von: Dong, Jing, et al.
Veröffentlicht: (2024)
The Limit Points of (Optimistic) Gradient Descent in Min-Max Optimization
von: Daskalakis, Constantinos, et al.
Veröffentlicht: (2018)
von: Daskalakis, Constantinos, et al.
Veröffentlicht: (2018)
Revisiting Character-level Adversarial Attacks for Language Models
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2024)
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2024)
Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games
von: Leonardos, Stefanos, et al.
Veröffentlicht: (2021)
von: Leonardos, Stefanos, et al.
Veröffentlicht: (2021)
Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
von: Xu, Hang, et al.
Veröffentlicht: (2024)
von: Xu, Hang, et al.
Veröffentlicht: (2024)
Gradient Manipulation in Distributed Stochastic Gradient Descent with Strategic Agents: Truthful Incentives with Convergence Guarantees
von: Chen, Ziqin, et al.
Veröffentlicht: (2026)
von: Chen, Ziqin, et al.
Veröffentlicht: (2026)
Optimistic Estimation of Convergence in Markov Chains with the Average-Mixing Time
von: Wolfer, Geoffrey, et al.
Veröffentlicht: (2024)
von: Wolfer, Geoffrey, et al.
Veröffentlicht: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
von: Zhou, Dongruo, et al.
Veröffentlicht: (2018)
von: Zhou, Dongruo, et al.
Veröffentlicht: (2018)
Optimistic Online Mirror Descent for Bridging Stochastic and Adversarial Online Convex Optimization
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Optimistic Online Learning in Symmetric Cone Games
von: Barakat, Anas, et al.
Veröffentlicht: (2025)
von: Barakat, Anas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Universal Gradient Methods for Stochastic Convex Optimization
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024) -
On the Generalization of Stochastic Gradient Descent with Momentum
von: Ramezani-Kebrya, Ali, et al.
Veröffentlicht: (2018) -
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
von: Viel, Stefano, et al.
Veröffentlicht: (2025) -
Polynomial Convergence of Bandit No-Regret Dynamics in Congestion Games
von: Dadi, Leello, et al.
Veröffentlicht: (2024) -
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)