A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Kihyuk, Tewari, Ambuj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
Offline Constrained Reinforcement Learning under Partial Data Coverage
von: Ko, Seokmin, et al.
Veröffentlicht: (2025)
von: Ko, Seokmin, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
von: Brooks, Marc, et al.
Veröffentlicht: (2025)
von: Brooks, Marc, et al.
Veröffentlicht: (2025)
Leveraging Offline Data in Linear Latent Contextual Bandits
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Controlling Statistical, Discretization, and Truncation Errors in Learning Fourier Linear Operators
von: Subedi, Unique, et al.
Veröffentlicht: (2024)
von: Subedi, Unique, et al.
Veröffentlicht: (2024)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
An Asymptotically Optimal Algorithm for the Convex Hull Membership Problem
von: Qiao, Gang, et al.
Veröffentlicht: (2023)
von: Qiao, Gang, et al.
Veröffentlicht: (2023)
Online Infinite-Dimensional Regression: Learning Linear Operators
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
Optimal Thresholding Linear Bandit
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2024)
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2024)
If generative AI is the answer, what is the question?
von: Tewari, Ambuj
Veröffentlicht: (2025)
von: Tewari, Ambuj
Veröffentlicht: (2025)
Operator Learning: A Statistical Perspective
von: Subedi, Unique, et al.
Veröffentlicht: (2025)
von: Subedi, Unique, et al.
Veröffentlicht: (2025)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
On the Benefits of Active Data Collection in Operator Learning
von: Subedi, Unique, et al.
Veröffentlicht: (2024)
von: Subedi, Unique, et al.
Veröffentlicht: (2024)
Learning to Partially Defer for Sequences
von: Rayan, Sahana, et al.
Veröffentlicht: (2025)
von: Rayan, Sahana, et al.
Veröffentlicht: (2025)
Is Zero-Shot Super-Resolution Possible in Operator Learning?
von: Subedi, Unique, et al.
Veröffentlicht: (2026)
von: Subedi, Unique, et al.
Veröffentlicht: (2026)
Quantum Learning Theory Beyond Batch Binary Classification
von: Mohan, Preetham, et al.
Veröffentlicht: (2023)
von: Mohan, Preetham, et al.
Veröffentlicht: (2023)
On the Minimax Regret in Online Ranking with Top-k Feedback
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2023)
Online Learning with Set-Valued Feedback
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
Distribution-Free Robust Predict-Then-Optimize in Function Spaces
von: Patel, Yash, et al.
Veröffentlicht: (2026)
von: Patel, Yash, et al.
Veröffentlicht: (2026)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
von: Mishra, Abhiti, et al.
Veröffentlicht: (2025)
von: Mishra, Abhiti, et al.
Veröffentlicht: (2025)
Operator Learning for Schrödinger Equation: Unitarity, Error Bounds, and Time Generalization
von: Patel, Yash, et al.
Veröffentlicht: (2025)
von: Patel, Yash, et al.
Veröffentlicht: (2025)
Compute Aligned Training: Optimizing for Test Time Inference
von: Ousherovitch, Adam, et al.
Veröffentlicht: (2026)
von: Ousherovitch, Adam, et al.
Veröffentlicht: (2026)
A Primal-Dual Algorithm for Hybrid Federated Learning
von: Overman, Tom, et al.
Veröffentlicht: (2022)
von: Overman, Tom, et al.
Veröffentlicht: (2022)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2026)
Sample Efficient Myopic Exploration Through Multitask Reinforcement Learning with Diverse Tasks
von: Xu, Ziping, et al.
Veröffentlicht: (2024)
von: Xu, Ziping, et al.
Veröffentlicht: (2024)
Near Optimal Pure Exploration in Logistic Bandits
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2024)
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2024)
Online Conformal Prediction: Enforcing monotonicity via Online Optimization
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2026)
von: Rivera, Eduardo Ochoa, et al.
Veröffentlicht: (2026)
Online Classification with Predictions
von: Raman, Vinod, et al.
Veröffentlicht: (2024)
von: Raman, Vinod, et al.
Veröffentlicht: (2024)
A Characterization of Multioutput Learnability
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
A Combinatorial Characterization of Supervised Online Learnability
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
von: Raman, Vinod, et al.
Veröffentlicht: (2023)
Joint Learning of Linear Time-Invariant Dynamical Systems
von: Modi, Aditya, et al.
Veröffentlicht: (2021)
von: Modi, Aditya, et al.
Veröffentlicht: (2021)
Off-Policy Primal-Dual Safe Reinforcement Learning
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025) -
Offline Constrained Reinforcement Learning under Partial Data Coverage
von: Ko, Seokmin, et al.
Veröffentlicht: (2025) -
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024) -
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024) -
Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
von: Brooks, Marc, et al.
Veröffentlicht: (2025)