Pausing Policy Learning in Non-stationary Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyunin, Jin, Ming, Lavaei, Javad, Sojoudi, Somayeh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
by: Ding, Yuhao, et al.
Published: (2021)
by: Ding, Yuhao, et al.
Published: (2021)
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
by: Ma, Ziye, et al.
Published: (2024)
by: Ma, Ziye, et al.
Published: (2024)
A CMDP-within-online framework for Meta-Safe Reinforcement Learning
by: Khattar, Vanshaj, et al.
Published: (2024)
by: Khattar, Vanshaj, et al.
Published: (2024)
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
by: Ying, Donghao, et al.
Published: (2022)
by: Ying, Donghao, et al.
Published: (2022)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
by: Li, Jingqi, et al.
Published: (2022)
by: Li, Jingqi, et al.
Published: (2022)
Distributionally Robust Joint Chance-Constrained Optimal Power Flow using Relative Entropy
by: Brock, Eli, et al.
Published: (2025)
by: Brock, Eli, et al.
Published: (2025)
Coordinating Distributed Energy Resources with Nodal Pricing in Distribution Networks: a Game-Theoretic Approach
by: Brock, Eli, et al.
Published: (2025)
by: Brock, Eli, et al.
Published: (2025)
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
GB-DQN: Gradient Boosted DQN Models for Non-stationary Reinforcement Learning
by: Lee, Chang-Hwan, et al.
Published: (2025)
by: Lee, Chang-Hwan, et al.
Published: (2025)
Transport of Algebraic Structure to Latent Embeddings
by: Pfrommer, Samuel, et al.
Published: (2024)
by: Pfrommer, Samuel, et al.
Published: (2024)
Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models
by: Lee, Hyunin, et al.
Published: (2025)
by: Lee, Hyunin, et al.
Published: (2025)
Safe Continual Reinforcement Learning in Non-stationary Environments
by: Coursey, Austin, et al.
Published: (2026)
by: Coursey, Austin, et al.
Published: (2026)
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Transformers Provably Learn to Internalize Chain-of-Thought
by: Huang, Yixiao, et al.
Published: (2026)
by: Huang, Yixiao, et al.
Published: (2026)
Subgradient Method for System Identification with Non-Smooth Objectives
by: Yalcin, Baturalp, et al.
Published: (2025)
by: Yalcin, Baturalp, et al.
Published: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
LLMs Should Express Uncertainty Explicitly
by: Guo, Junyu, et al.
Published: (2026)
by: Guo, Junyu, et al.
Published: (2026)
Forecasting in Offline Reinforcement Learning for Non-stationary Environments
by: Ada, Suzan Ece, et al.
Published: (2025)
by: Ada, Suzan Ece, et al.
Published: (2025)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Optimization Solution Functions as Deterministic Policies for Offline Reinforcement Learning
by: Khattar, Vanshaj, et al.
Published: (2024)
by: Khattar, Vanshaj, et al.
Published: (2024)
Mixing Classifiers to Alleviate the Accuracy-Robustness Trade-Off
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
by: Anderson, Brendon G., et al.
Published: (2021)
by: Anderson, Brendon G., et al.
Published: (2021)
MetaCURL: Non-stationary Concave Utility Reinforcement Learning
by: Moreno, Bianca Marin, et al.
Published: (2024)
by: Moreno, Bianca Marin, et al.
Published: (2024)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
by: Cheng, Ziheng, et al.
Published: (2026)
by: Cheng, Ziheng, et al.
Published: (2026)
Non-stationary and Varying-discounting Markov Decision Processes for Reinforcement Learning
by: Chen, Zhizuo, et al.
Published: (2025)
by: Chen, Zhizuo, et al.
Published: (2025)
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
by: Bai, Yatong, et al.
Published: (2025)
by: Bai, Yatong, et al.
Published: (2025)
Parameter Stress Analysis in Reinforcement Learning: Applying Synaptic Filtering to Policy Networks
by: Abdeen, Zain ul, et al.
Published: (2025)
by: Abdeen, Zain ul, et al.
Published: (2025)
Efficient Methods for Non-stationary Online Learning
by: Zhao, Peng, et al.
Published: (2023)
by: Zhao, Peng, et al.
Published: (2023)
Efficient Global Optimization of Two-Layer ReLU Networks: Quadratic-Time Algorithms and Adversarial Training
by: Bai, Yatong, et al.
Published: (2022)
by: Bai, Yatong, et al.
Published: (2022)
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
A theory on the absence of spurious solutions for nonconvex and nonsmooth optimization
by: Josz, Cedric, et al.
Published: (2018)
by: Josz, Cedric, et al.
Published: (2018)
High Probability Complexity Bounds of Trust-Region Stochastic Sequential Quadratic Programming with Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2025)
by: Fang, Yuchen, et al.
Published: (2025)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization
by: Akgül, Abdullah, et al.
Published: (2025)
by: Akgül, Abdullah, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Similar Items
-
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
by: Ding, Yuhao, et al.
Published: (2021) -
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025) -
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
by: Ma, Ziye, et al.
Published: (2024) -
A CMDP-within-online framework for Meta-Safe Reinforcement Learning
by: Khattar, Vanshaj, et al.
Published: (2024) -
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
by: Ying, Donghao, et al.
Published: (2022)