A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zeng, Sihan, Bhatt, Sujay, Ganesh, Sumitra, Koppel, Alec |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning in Herding Mean Field Games: Single-Loop Algorithm with Finite-Time Convergence Analysis
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
Partially Observable Contextual Bandits with Linear Payoffs
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
Regularized Proportional Fairness Mechanism for Resource Allocation Without Money
por: Zeng, Sihan, et al.
Publicado: (2025)
por: Zeng, Sihan, et al.
Publicado: (2025)
Natural Policy Gradient and Actor Critic Methods for Constrained Multi-Task Reinforcement Learning
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
Learning Payment-Free Resource Allocation Mechanisms
por: Zeng, Sihan, et al.
Publicado: (2023)
por: Zeng, Sihan, et al.
Publicado: (2023)
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
por: Zeng, Sihan, et al.
Publicado: (2021)
por: Zeng, Sihan, et al.
Publicado: (2021)
Rethinking Neural Network Learning Rates: A Stackelberg Perspective
por: Zeng, Sihan, et al.
Publicado: (2026)
por: Zeng, Sihan, et al.
Publicado: (2026)
Learning in Stackelberg Mean Field Games: A Non-Asymptotic Analysis
por: Zeng, Sihan, et al.
Publicado: (2025)
por: Zeng, Sihan, et al.
Publicado: (2025)
Approximate Equivariance in Reinforcement Learning
por: Park, Jung Yeon, et al.
Publicado: (2024)
por: Park, Jung Yeon, et al.
Publicado: (2024)
Constrained Bi-Level Optimization: Proximal Lagrangian Value function Approach and Hessian-free Algorithm
por: Yao, Wei, et al.
Publicado: (2024)
por: Yao, Wei, et al.
Publicado: (2024)
Fast Two-Time-Scale Stochastic Gradient Method with Applications in Reinforcement Learning
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
A Two-Time-Scale Stochastic Optimization Framework with Applications in Control and Reinforcement Learning
por: Zeng, Sihan, et al.
Publicado: (2021)
por: Zeng, Sihan, et al.
Publicado: (2021)
Moreau Envelope for Nonconvex Bi-Level Optimization: A Single-loop and Hessian-free Solution Strategy
por: Liu, Risheng, et al.
Publicado: (2024)
por: Liu, Risheng, et al.
Publicado: (2024)
Sharpened Lazy Incremental Quasi-Newton Method
por: Lahoti, Aakash, et al.
Publicado: (2023)
por: Lahoti, Aakash, et al.
Publicado: (2023)
A Communication-Efficient Decentralized Actor-Critic Algorithm
por: Ren, Xiaoxing, et al.
Publicado: (2025)
por: Ren, Xiaoxing, et al.
Publicado: (2025)
Weak Convergence Analysis of Online Neural Actor-Critic Algorithms
por: Lam, Samuel Chun-Hei, et al.
Publicado: (2024)
por: Lam, Samuel Chun-Hei, et al.
Publicado: (2024)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
por: Cui, Mingxuan, et al.
Publicado: (2025)
por: Cui, Mingxuan, et al.
Publicado: (2025)
UVIP: Model-Free Approach to Evaluate Reinforcement Learning Algorithms
por: Belomestny, Denis, et al.
Publicado: (2021)
por: Belomestny, Denis, et al.
Publicado: (2021)
Quasi-Newton Compatible Actor-Critic for Deterministic Policies
por: Kordabad, Arash Bahari, et al.
Publicado: (2025)
por: Kordabad, Arash Bahari, et al.
Publicado: (2025)
Optimal Hessian/Jacobian-Free Nonconvex-PL Bilevel Optimization
por: Huang, Feihu
Publicado: (2024)
por: Huang, Feihu
Publicado: (2024)
QCQP-Net: Reliably Learning Feasible Alternating Current Optimal Power Flow Solutions Under Constraints
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
por: Zhang, Yufeng, et al.
Publicado: (2021)
por: Zhang, Yufeng, et al.
Publicado: (2021)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
por: Parfenov, Valery, et al.
Publicado: (2026)
por: Parfenov, Valery, et al.
Publicado: (2026)
Control Theoretic Approach to Fine-Tuning and Transfer Learning
por: Bayram, Erkan, et al.
Publicado: (2024)
por: Bayram, Erkan, et al.
Publicado: (2024)
CACTO-SL: Using Sobolev Learning to improve Continuous Actor-Critic with Trajectory Optimization
por: Alboni, Elisa, et al.
Publicado: (2023)
por: Alboni, Elisa, et al.
Publicado: (2023)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
por: Hamza, Ishaq, et al.
Publicado: (2026)
por: Hamza, Ishaq, et al.
Publicado: (2026)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
por: Zhao, Hanyang, et al.
Publicado: (2025)
por: Zhao, Hanyang, et al.
Publicado: (2025)
Stochastic Hessian Fittings with Lie Groups
por: Li, Xi-Lin
Publicado: (2024)
por: Li, Xi-Lin
Publicado: (2024)
Tuning-Free Stochastic Optimization
por: Khaled, Ahmed, et al.
Publicado: (2024)
por: Khaled, Ahmed, et al.
Publicado: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
por: Kharrat, Salma, et al.
Publicado: (2024)
por: Kharrat, Salma, et al.
Publicado: (2024)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
por: Semenov, Andrei, et al.
Publicado: (2025)
por: Semenov, Andrei, et al.
Publicado: (2025)
Towards Quantifying the Hessian Structure of Neural Networks
por: Dong, Zhaorui, et al.
Publicado: (2025)
por: Dong, Zhaorui, et al.
Publicado: (2025)
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
por: Yang, Tong, et al.
Publicado: (2025)
por: Yang, Tong, et al.
Publicado: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
por: Sun, Ying, et al.
Publicado: (2024)
por: Sun, Ying, et al.
Publicado: (2024)
Extensions of Robbins-Siegmund Theorem with Applications in Reinforcement Learning
por: Liu, Xinyu, et al.
Publicado: (2025)
por: Liu, Xinyu, et al.
Publicado: (2025)
Unlocking TriLevel Learning with Level-Wise Zeroth Order Constraints: Distributed Algorithms and Provable Non-Asymptotic Convergence
por: Jiao, Yang, et al.
Publicado: (2024)
por: Jiao, Yang, et al.
Publicado: (2024)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
por: Bylinkin, Dmitry, et al.
Publicado: (2024)
por: Bylinkin, Dmitry, et al.
Publicado: (2024)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Mitigating Covariate Shift in Misspecified Regression with Applications to Reinforcement Learning
por: Amortila, Philip, et al.
Publicado: (2024)
por: Amortila, Philip, et al.
Publicado: (2024)
Ejemplares similares
-
Learning in Herding Mean Field Games: Single-Loop Algorithm with Finite-Time Convergence Analysis
por: Zeng, Sihan, et al.
Publicado: (2024) -
Partially Observable Contextual Bandits with Linear Payoffs
por: Zeng, Sihan, et al.
Publicado: (2024) -
Regularized Proportional Fairness Mechanism for Resource Allocation Without Money
por: Zeng, Sihan, et al.
Publicado: (2025) -
Natural Policy Gradient and Actor Critic Methods for Constrained Multi-Task Reinforcement Learning
por: Zeng, Sihan, et al.
Publicado: (2024) -
Learning Payment-Free Resource Allocation Mechanisms
por: Zeng, Sihan, et al.
Publicado: (2023)