Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Han, Yang, Zhuoran, Chen, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Penalty-based Bilevel Gradient Descent Method
von: Shen, Han, et al.
Veröffentlicht: (2023)
von: Shen, Han, et al.
Veröffentlicht: (2023)
Efficient Penalty-Based Bilevel Methods: Improved Analysis, Novel Updates, and Flatness Condition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
A Primal-Dual-Assisted Penalty Approach to Bilevel Optimization with Coupled Constraints
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2024)
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
On Penalty Methods for Nonconvex Bilevel Optimization and First-Order Stochastic Approximation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2023)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2023)
Unlocking Global Optimality in Bilevel Optimization: A Pilot Study
von: Xiao, Quan, et al.
Veröffentlicht: (2024)
von: Xiao, Quan, et al.
Veröffentlicht: (2024)
A Penalty-Based Method for Communication-Efficient Decentralized Bilevel Programming
von: Nazari, Parvin, et al.
Veröffentlicht: (2022)
von: Nazari, Parvin, et al.
Veröffentlicht: (2022)
Accelerated Gradient Methods for Sparse Statistical Learning with Nonconvex Penalties
von: Yang, Kai, et al.
Veröffentlicht: (2020)
von: Yang, Kai, et al.
Veröffentlicht: (2020)
Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency
von: Cai, Qi, et al.
Veröffentlicht: (2022)
von: Cai, Qi, et al.
Veröffentlicht: (2022)
Faster Gradient Methods for Highly-Smooth Stochastic Bilevel Optimization
von: Chen, Lesi, et al.
Veröffentlicht: (2025)
von: Chen, Lesi, et al.
Veröffentlicht: (2025)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
von: Thoma, Vinzenz, et al.
Veröffentlicht: (2024)
von: Thoma, Vinzenz, et al.
Veröffentlicht: (2024)
A First-order Generative Bilevel Optimization Framework for Diffusion Models
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
An Adaptively Inexact Method for Bilevel Learning Using Primal-Dual Style Differentiation
von: Bogensperger, Lea, et al.
Veröffentlicht: (2024)
von: Bogensperger, Lea, et al.
Veröffentlicht: (2024)
On the Complexity of First-Order Methods in Stochastic Bilevel Optimization
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
First-Order Methods for Linearly Constrained Bilevel Optimization
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Bilevel Learning with Inexact Stochastic Gradients
von: Salehi, Mohammad Sadegh, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammad Sadegh, et al.
Veröffentlicht: (2024)
An Accelerated Gradient Method for Convex Smooth Simple Bilevel Optimization
von: Cao, Jincheng, et al.
Veröffentlicht: (2024)
von: Cao, Jincheng, et al.
Veröffentlicht: (2024)
Accelerated Fully First-Order Methods for Bilevel and Minimax Optimization
von: Li, Chris Junchi
Veröffentlicht: (2024)
von: Li, Chris Junchi
Veröffentlicht: (2024)
A Framework for Bilevel Optimization on Riemannian Manifolds
von: Han, Andi, et al.
Veröffentlicht: (2024)
von: Han, Andi, et al.
Veröffentlicht: (2024)
Penalty-based Methods for Simple Bilevel Optimization under Hölderian Error Bounds
von: Chen, Pengyu, et al.
Veröffentlicht: (2024)
von: Chen, Pengyu, et al.
Veröffentlicht: (2024)
Bilevel Models for Adversarial Learning and A Case Study
von: Zheng, Yutong, et al.
Veröffentlicht: (2025)
von: Zheng, Yutong, et al.
Veröffentlicht: (2025)
FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
von: Chen, Lisha, et al.
Veröffentlicht: (2024)
von: Chen, Lisha, et al.
Veröffentlicht: (2024)
Position: Adopt Constraints Over Fixed Penalties in Deep Learning
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
von: Ramirez, Juan, et al.
Veröffentlicht: (2025)
Self-Supervised Penalty-Based Learning for Robust Constrained Optimization
von: Benslimane, Wyame, et al.
Veröffentlicht: (2025)
von: Benslimane, Wyame, et al.
Veröffentlicht: (2025)
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
von: Zhang, Yufeng, et al.
Veröffentlicht: (2020)
von: Zhang, Yufeng, et al.
Veröffentlicht: (2020)
Fully First-Order Algorithms for Online Bilevel Optimization
von: Jia, Tingkai, et al.
Veröffentlicht: (2026)
von: Jia, Tingkai, et al.
Veröffentlicht: (2026)
Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach
von: Fernando, Heshan, et al.
Veröffentlicht: (2022)
von: Fernando, Heshan, et al.
Veröffentlicht: (2022)
Learning to Solve Constrained Bilevel Control Co-Design Problems
von: Kotary, James, et al.
Veröffentlicht: (2025)
von: Kotary, James, et al.
Veröffentlicht: (2025)
Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization
von: Dong, Youran, et al.
Veröffentlicht: (2025)
von: Dong, Youran, et al.
Veröffentlicht: (2025)
Bayesian Optimization of Bilevel Problems
von: Ekmekcioglu, Omer, et al.
Veröffentlicht: (2024)
von: Ekmekcioglu, Omer, et al.
Veröffentlicht: (2024)
On Stability in Optimistic Bilevel Optimization
von: Royset, Johannes O.
Veröffentlicht: (2024)
von: Royset, Johannes O.
Veröffentlicht: (2024)
Proximal Operators of Sorted Nonconvex Penalties
von: Gagneux, Anne, et al.
Veröffentlicht: (2025)
von: Gagneux, Anne, et al.
Veröffentlicht: (2025)
Functionally Constrained Algorithm Solves Convex Simple Bilevel Problems
von: Zhang, Huaqing, et al.
Veröffentlicht: (2024)
von: Zhang, Huaqing, et al.
Veröffentlicht: (2024)
Online Bilevel Optimization: Regret Analysis of Online Alternating Gradient Methods
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2022)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2022)
Safe Gradient Flow for Bilevel Optimization
von: Sharifi, Sina, et al.
Veröffentlicht: (2025)
von: Sharifi, Sina, et al.
Veröffentlicht: (2025)
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
von: Zhang, Yufeng, et al.
Veröffentlicht: (2021)
von: Zhang, Yufeng, et al.
Veröffentlicht: (2021)
Adaptive Learning-based Surrogate Method for Stochastic Programs with Implicitly Decision-dependent Uncertainty
von: Shen, Boyang, et al.
Veröffentlicht: (2025)
von: Shen, Boyang, et al.
Veröffentlicht: (2025)
Online Nonconvex Bilevel Optimization with Bregman Divergences
von: Bohne, Jason, et al.
Veröffentlicht: (2024)
von: Bohne, Jason, et al.
Veröffentlicht: (2024)
Problem-Parameter-Free Decentralized Bilevel Optimization
von: Zhai, Zhiwei, et al.
Veröffentlicht: (2025)
von: Zhai, Zhiwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Penalty-based Bilevel Gradient Descent Method
von: Shen, Han, et al.
Veröffentlicht: (2023) -
Efficient Penalty-Based Bilevel Methods: Improved Analysis, Novel Updates, and Flatness Condition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025) -
A Primal-Dual-Assisted Penalty Approach to Bilevel Optimization with Coupled Constraints
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2024) -
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
von: Shen, Yiyang, et al.
Veröffentlicht: (2026) -
On Penalty Methods for Nonconvex Bilevel Optimization and First-Order Stochastic Approximation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2023)