Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Kitamura, Toshinori, Ghosh, Arnob, Kozuno, Tadashi, Kumagai, Wataru, Kasaura, Kazumi, Hoshino, Kenta, Hosoe, Yohei, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
by: Nakao, Tomoki, et al.
Published: (2025)
by: Nakao, Tomoki, et al.
Published: (2025)
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
by: Roknilamouki, Amirhossein, et al.
Published: (2025)
by: Roknilamouki, Amirhossein, et al.
Published: (2025)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
by: Kitamura, Toshinori, et al.
Published: (2026)
by: Kitamura, Toshinori, et al.
Published: (2026)
Compactness in Constructive Mathematics via Affine Logic
by: Kasaura, Kazumi
Published: (2026)
by: Kasaura, Kazumi
Published: (2026)
Homotopy-Aware Multi-Agent Path Planning on Plane
by: Kasaura, Kazumi
Published: (2023)
by: Kasaura, Kazumi
Published: (2023)
Generation of Geodesics with Actor-Critic Reinforcement Learning to Predict Midpoints
by: Kasaura, Kazumi
Published: (2024)
by: Kasaura, Kazumi
Published: (2024)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Stochastic Safety-critical Control Compensating Safety Probability for Marine Vessel Tracking
by: Matsuo, Too, et al.
Published: (2026)
by: Matsuo, Too, et al.
Published: (2026)
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Control Barrier Functions for Stochastic Systems and Safety-critical Control Designs
by: Nishimura, Yuki, et al.
Published: (2022)
by: Nishimura, Yuki, et al.
Published: (2022)
Self Iterative Label Refinement via Robust Unlabeled Learning
by: Asano, Hikaru, et al.
Published: (2025)
by: Asano, Hikaru, et al.
Published: (2025)
Multi-Agent Behavior Retrieval: Retrieval-Augmented Policy Training for Cooperative Push Manipulation by Mobile Robots
by: Kuroki, So, et al.
Published: (2023)
by: Kuroki, So, et al.
Published: (2023)
Uniformly boundedness of finite Morse index solutions to semilinear elliptic equations with rapidly growing nonlinearities in two dimensions
by: Kumagai, Kenta
Published: (2025)
by: Kumagai, Kenta
Published: (2025)
Bifurcation diagrams for semilinear elliptic equations with singular weights in two dimensions
by: Kumagai, Kenta
Published: (2024)
by: Kumagai, Kenta
Published: (2024)
Bifurcation diagrams of semilinear elliptic equations for supercritical nonlinearities in two dimensions
by: Kumagai, Kenta
Published: (2024)
by: Kumagai, Kenta
Published: (2024)
Classification of bifurcation structure for semilinear elliptic equations in a ball
by: Kumagai, Kenta
Published: (2025)
by: Kumagai, Kenta
Published: (2025)
Assessing Human Intelligence Augmentation Strategies Using Brain Machine Interfaces and Brain Organoids in the Era of AI Advancement
by: Kitamura, Kenta
Published: (2025)
by: Kitamura, Kenta
Published: (2025)
Physics-informed RL for Maximal Safety Probability Estimation
by: Hoshino, Hikaru, et al.
Published: (2024)
by: Hoshino, Hikaru, et al.
Published: (2024)
A Linear-Time Algorithm for the Closest Vector Problem of Triangular Lattices
by: Takahashi, Kenta, et al.
Published: (2024)
by: Takahashi, Kenta, et al.
Published: (2024)
Lean Formalization of Generalization Error Bound by Rademacher Complexity and Dudley's Entropy Integral
by: Sonoda, Sho, et al.
Published: (2025)
by: Sonoda, Sho, et al.
Published: (2025)
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
by: Asano, Hikaru, et al.
Published: (2026)
by: Asano, Hikaru, et al.
Published: (2026)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
When to Replan? An Adaptive Replanning Strategy for Autonomous Navigation using Deep Reinforcement Learning
by: Honda, Kohei, et al.
Published: (2023)
by: Honda, Kohei, et al.
Published: (2023)
A note on quantum subgroups of free quantum groups
by: Hoshino, Mao, et al.
Published: (2024)
by: Hoshino, Mao, et al.
Published: (2024)
A Multi‐Objective Evolutionary Algorithm Based on Bilayered Decomposition for Constrained Multi‐Objective Optimization
by: Yusuke Yasuda, et al.
Published: (2024)
by: Yusuke Yasuda, et al.
Published: (2024)
Safety-Critical Control for Discrete-time Stochastic Systems with Flexible Safe Bounds using Affine and Quadratic Control Barrier Functions
by: Fushimi, Sotaro, et al.
Published: (2025)
by: Fushimi, Sotaro, et al.
Published: (2025)
Comments on resolution of nonassociativity in SFT- an example from axioms of BCFT-
by: Matsuo Yutaka
Published: (2002)
by: Matsuo Yutaka
Published: (2002)
Singular solutions and bifurcation diagram of semilinear elliptic equations with general nonlinearity in two dimensions
by: Kikuchi, Hiroaki, et al.
Published: (2025)
by: Kikuchi, Hiroaki, et al.
Published: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
by: Zuo, Qian, et al.
Published: (2025)
by: Zuo, Qian, et al.
Published: (2025)
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
by: Kasaura, Kazumi, et al.
Published: (2025)
by: Kasaura, Kazumi, et al.
Published: (2025)
LeanConjecturer: Automatic Generation of Mathematical Conjectures for Theorem Proving
by: Onda, Naoto, et al.
Published: (2025)
by: Onda, Naoto, et al.
Published: (2025)
Learning Kernel-Based MDPs from Episodic Preferential Feedback
by: Pavlovic, Nikola, et al.
Published: (2026)
by: Pavlovic, Nikola, et al.
Published: (2026)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
by: Dong, Zixuan, et al.
Published: (2022)
by: Dong, Zixuan, et al.
Published: (2022)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
by: Liu, Xingtu, et al.
Published: (2025)
by: Liu, Xingtu, et al.
Published: (2025)
Similar Items
-
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
by: Kitamura, Toshinori, et al.
Published: (2024) -
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
by: Kitamura, Toshinori, et al.
Published: (2024) -
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
by: Nakao, Tomoki, et al.
Published: (2025) -
Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
by: Roknilamouki, Amirhossein, et al.
Published: (2025) -
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)