Offline RL via Feature-Occupancy Gradient Ascent
Fuente:
arXiv
Saved in:
| Main Authors: | Neu, Gergely, Okolo, Nneka |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dealing with unbounded gradients in stochastic saddle-point optimization
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
Online-to-PAC Conversions: Generalization Bounds via Regret Analysis
by: Lugosi, Gábor, et al.
Published: (2023)
by: Lugosi, Gábor, et al.
Published: (2023)
Online combinatorial optimization with stochastic decision sets and adversarial losses
by: Neu, Gergely, et al.
Published: (2026)
by: Neu, Gergely, et al.
Published: (2026)
Reinforcement Learning in POMDP's via Direct Gradient Ascent
by: Baxter, Jonathan, et al.
Published: (2025)
by: Baxter, Jonathan, et al.
Published: (2025)
Generalization bounds for mixing processes via delayed online-to-PAC conversions
by: Abeles, Baptiste, et al.
Published: (2024)
by: Abeles, Baptiste, et al.
Published: (2024)
Online-to-PAC generalization bounds under graph-mixing dependencies
by: Abélès, Baptiste, et al.
Published: (2024)
by: Abélès, Baptiste, et al.
Published: (2024)
Optimistic Information Directed Sampling
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
Online learning with Erdős-Rényi side-observation graphs
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Online learning with noisy side observations
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Sparse Optimistic Information Directed Sampling
by: Schwartz, Ludovic, et al.
Published: (2025)
by: Schwartz, Ludovic, et al.
Published: (2025)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
by: Moulin, Antoine, et al.
Published: (2025)
by: Moulin, Antoine, et al.
Published: (2025)
Posterior Approximation using Stochastic Gradient Ascent with Adaptive Stepsize
by: Lim, Kart-Leong, et al.
Published: (2024)
by: Lim, Kart-Leong, et al.
Published: (2024)
Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
by: Huang, Zhuo, et al.
Published: (2026)
by: Huang, Zhuo, et al.
Published: (2026)
MoMA: Model-based Mirror Ascent for Offline Reinforcement Learning
by: Hong, Mao, et al.
Published: (2024)
by: Hong, Mao, et al.
Published: (2024)
Confidence Sequences for Generalized Linear Models via Regret Analysis
by: Clerico, Eugenio, et al.
Published: (2025)
by: Clerico, Eugenio, et al.
Published: (2025)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
On Gradient Descent Ascent for Nonconvex-Concave Minimax Problems
by: Lin, Tianyi, et al.
Published: (2019)
by: Lin, Tianyi, et al.
Published: (2019)
Linear Bandits with Non-i.i.d. Noise
by: Abélès, Baptiste, et al.
Published: (2025)
by: Abélès, Baptiste, et al.
Published: (2025)
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
by: Zhang, Qixin, et al.
Published: (2024)
by: Zhang, Qixin, et al.
Published: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
by: Mukherjee, Subhojyoti, et al.
Published: (2026)
Offline Imitation from Observation via Primal Wasserstein State Occupancy Matching
by: Yan, Kai, et al.
Published: (2023)
by: Yan, Kai, et al.
Published: (2023)
Improving Offline RL by Blending Heuristics
by: Geng, Sinong, et al.
Published: (2023)
by: Geng, Sinong, et al.
Published: (2023)
Negative Stepsizes Make Gradient-Descent-Ascent Converge
by: Shugart, Henry, et al.
Published: (2025)
by: Shugart, Henry, et al.
Published: (2025)
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
by: Lin, Tianyi, et al.
Published: (2024)
by: Lin, Tianyi, et al.
Published: (2024)
Sobolev Gradient Ascent for Optimal Transport: Barycenter Optimization and Convergence Analysis
by: Kim, Kaheon, et al.
Published: (2025)
by: Kim, Kaheon, et al.
Published: (2025)
Decentralized Stochastic Gradient Descent Ascent for Finite-Sum Minimax Problems
by: Gao, Hongchang
Published: (2022)
by: Gao, Hongchang
Published: (2022)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
From Robotics to Sepsis Treatment: Offline RL via Geometric Pessimism
by: Wanjari, Sarthak
Published: (2026)
by: Wanjari, Sarthak
Published: (2026)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
by: Shen, Wei, et al.
Published: (2023)
by: Shen, Wei, et al.
Published: (2023)
When Are RL Hyperparameters Benign? A Study in Offline Goal-Conditioned RL
by: Töpperwien, Jan Malte, et al.
Published: (2026)
by: Töpperwien, Jan Malte, et al.
Published: (2026)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
by: Beck, Jacob
Published: (2025)
by: Beck, Jacob
Published: (2025)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)
by: Wang, Qi, et al.
Published: (2023)
Augmenting Offline RL with Unlabeled Data
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Selective Uncertainty Propagation in Offline RL
by: Krishnamurthy, Sanath Kumar, et al.
Published: (2023)
by: Krishnamurthy, Sanath Kumar, et al.
Published: (2023)
Decoupled Prioritized Resampling for Offline RL
by: Yue, Yang, et al.
Published: (2023)
by: Yue, Yang, et al.
Published: (2023)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
Similar Items
-
Dealing with unbounded gradients in stochastic saddle-point optimization
by: Neu, Gergely, et al.
Published: (2024) -
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026) -
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
by: Moulin, Antoine, et al.
Published: (2025) -
Online-to-PAC Conversions: Generalization Bounds via Regret Analysis
by: Lugosi, Gábor, et al.
Published: (2023) -
Online combinatorial optimization with stochastic decision sets and adversarial losses
by: Neu, Gergely, et al.
Published: (2026)