Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pavse, Brahma S., Zurek, Matthew, Chen, Yudong, Xie, Qiaomin, Hanna, Josiah P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stable Offline Value Function Learning with Bisimulation-based Representations
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
Faster Fixed-Point Methods for Multichain MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Prelimit Coupling and Steady-State Convergence of Constant-stepsize Nonsmooth Contractive SA
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Gap-Free Clustering: Sensitivity and Robustness of SDP
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
Contextual Online Pricing with (Biased) Offline Data
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
von: Huo, Dongyan, et al.
Veröffentlicht: (2024)
von: Huo, Dongyan, et al.
Veröffentlicht: (2024)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
A Piecewise Lyapunov Analysis of Sub-quadratic SGD: Applications to Robust and Quantile Regression
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Optimal Attack and Defense for Reinforcement Learning
von: McMahan, Jeremy, et al.
Veröffentlicht: (2023)
von: McMahan, Jeremy, et al.
Veröffentlicht: (2023)
Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces
von: Flynn, Hamish, et al.
Veröffentlicht: (2026)
von: Flynn, Hamish, et al.
Veröffentlicht: (2026)
Reinforcement Learning via Auxiliary Task Distillation
von: Harish, Abhinav Narayan, et al.
Veröffentlicht: (2024)
von: Harish, Abhinav Narayan, et al.
Veröffentlicht: (2024)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
von: Hong, Yige, et al.
Veröffentlicht: (2023)
von: Hong, Yige, et al.
Veröffentlicht: (2023)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
von: Hong, Yige, et al.
Veröffentlicht: (2024)
von: Hong, Yige, et al.
Veröffentlicht: (2024)
Wasserstein-p Central Limit Theorem Rates: From Local Dependence to Markov Chains
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
Optimally Installing Strict Equilibria
von: McMahan, Jeremy, et al.
Veröffentlicht: (2025)
von: McMahan, Jeremy, et al.
Veröffentlicht: (2025)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Inception: Efficiently Computable Misinformation Attacks on Markov Games
von: McMahan, Jeremy, et al.
Veröffentlicht: (2024)
von: McMahan, Jeremy, et al.
Veröffentlicht: (2024)
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Control of Non-Markovian Cellular Population Dynamics
von: Kratz, Josiah C., et al.
Veröffentlicht: (2024)
von: Kratz, Josiah C., et al.
Veröffentlicht: (2024)
The Limits of Transfer Reinforcement Learning with Latent Low-rank Structure
von: Sam, Tyler, et al.
Veröffentlicht: (2024)
von: Sam, Tyler, et al.
Veröffentlicht: (2024)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
von: Hong, Yige, et al.
Veröffentlicht: (2024)
von: Hong, Yige, et al.
Veröffentlicht: (2024)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Active Advantage-Aligned Online Reinforcement Learning with Offline Data
von: Liu, Xuefeng, et al.
Veröffentlicht: (2025)
von: Liu, Xuefeng, et al.
Veröffentlicht: (2025)
Harnessing Density Ratios for Online Reinforcement Learning
von: Amortila, Philip, et al.
Veröffentlicht: (2024)
von: Amortila, Philip, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Stable Offline Value Function Learning with Bisimulation-based Representations
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024) -
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024) -
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023) -
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
von: Huo, Dongyan, et al.
Veröffentlicht: (2022) -
Faster Fixed-Point Methods for Multichain MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)