End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mhammedi, Zakaria, Rakhlin, Alexander, Okolo, Nneka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
von: Mhammedi, Zakaria
Veröffentlicht: (2024)
von: Mhammedi, Zakaria
Veröffentlicht: (2024)
Efficient Model-Free Exploration in Low-Rank MDPs
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2023)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2023)
Offline RL via Feature-Occupancy Gradient Ascent
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
The Power of Resets in Online Reinforcement Learning
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
Online Convex Optimization with a Separation Oracle
von: Mhammedi, Zakaria
Veröffentlicht: (2024)
von: Mhammedi, Zakaria
Veröffentlicht: (2024)
Dealing with unbounded gradients in stochastic saddle-point optimization
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
von: Neu, Gergely, et al.
Veröffentlicht: (2024)
Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2026)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2026)
Fully Unconstrained Online Learning
von: Cutkosky, Ashok, et al.
Veröffentlicht: (2024)
von: Cutkosky, Ashok, et al.
Veröffentlicht: (2024)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
von: Foster, Dylan J., et al.
Veröffentlicht: (2025)
von: Foster, Dylan J., et al.
Veröffentlicht: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
von: Chen, Fan, et al.
Veröffentlicht: (2024)
von: Chen, Fan, et al.
Veröffentlicht: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
von: Golowich, Noah, et al.
Veröffentlicht: (2024)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Adaptive Matrix Online Learning through Smoothing with Guarantees for Nonsmooth Nonconvex Optimization
von: Jiang, Ruichen, et al.
Veröffentlicht: (2026)
von: Jiang, Ruichen, et al.
Veröffentlicht: (2026)
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2025)
von: Roknilamouki, Amirhossein, et al.
Veröffentlicht: (2025)
Deterministic Exploration via Stationary Bellman Error Maximization
von: Griesbach, Sebastian, et al.
Veröffentlicht: (2024)
von: Griesbach, Sebastian, et al.
Veröffentlicht: (2024)
Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
von: Khurram, Aleesha, et al.
Veröffentlicht: (2025)
von: Khurram, Aleesha, et al.
Veröffentlicht: (2025)
One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
von: Chen, Elynn, et al.
Veröffentlicht: (2026)
Oracle-Efficient Smoothed Online Learning for Piecewise Continuous Decision Making
von: Block, Adam, et al.
Veröffentlicht: (2023)
von: Block, Adam, et al.
Veröffentlicht: (2023)
FullCert: Deterministic End-to-End Certification for Training and Inference of Neural Networks
von: Lorenz, Tobias, et al.
Veröffentlicht: (2024)
von: Lorenz, Tobias, et al.
Veröffentlicht: (2024)
Hierarchical End-to-End Taylor Bounds for Complete Neural Network Verification
von: Entesari, Taha, et al.
Veröffentlicht: (2026)
von: Entesari, Taha, et al.
Veröffentlicht: (2026)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
von: Lyu, Lixing, et al.
Veröffentlicht: (2025)
von: Lyu, Lixing, et al.
Veröffentlicht: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States
von: Chen, Yujiao
Veröffentlicht: (2026)
von: Chen, Yujiao
Veröffentlicht: (2026)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
End-to-end RL Improves Dexterous Grasping Policies
von: Singh, Ritvik, et al.
Veröffentlicht: (2025)
von: Singh, Ritvik, et al.
Veröffentlicht: (2025)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
von: Muni, Aneri, et al.
Veröffentlicht: (2026)
von: Muni, Aneri, et al.
Veröffentlicht: (2026)
Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
von: Chen, Fan, et al.
Veröffentlicht: (2026)
von: Chen, Fan, et al.
Veröffentlicht: (2026)
Efficient End-to-End Learning for Decision-Making: A Meta-Optimization Approach
von: Cristian, Rares, et al.
Veröffentlicht: (2025)
von: Cristian, Rares, et al.
Veröffentlicht: (2025)
Solving robust MDPs as a sequence of static RL problems
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
von: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Veröffentlicht: (2025)
von: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Veröffentlicht: (2025)
RANDPOL: Parameter-Efficient End-to-End Quadruped Locomotion via Randomized Policy Learning
von: Liu, Zhuochen, et al.
Veröffentlicht: (2025)
von: Liu, Zhuochen, et al.
Veröffentlicht: (2025)
Coordinate Descent for Network Linearization
von: Rakhlin, Vlad, et al.
Veröffentlicht: (2025)
von: Rakhlin, Vlad, et al.
Veröffentlicht: (2025)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
von: Amortila, Philip, et al.
Veröffentlicht: (2024)
von: Amortila, Philip, et al.
Veröffentlicht: (2024)
Deep Riemannian Networks for End-to-End EEG Decoding
von: Wilson, Daniel, et al.
Veröffentlicht: (2022)
von: Wilson, Daniel, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
von: Mhammedi, Zakaria
Veröffentlicht: (2024) -
Efficient Model-Free Exploration in Low-Rank MDPs
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2023) -
Offline RL via Feature-Occupancy Gradient Ascent
von: Neu, Gergely, et al.
Veröffentlicht: (2024) -
The Power of Resets in Online Reinforcement Learning
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024) -
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)