A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kitamura, Toshinori, Kozuno, Tadashi, Kato, Masahiro, Ichihara, Yuki, Nishimori, Soichiro, Sannai, Akiyoshi, Sonoda, Sho, Kumagai, Wataru, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2026)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2026)
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
von: Kasaura, Kazumi, et al.
Veröffentlicht: (2025)
von: Kasaura, Kazumi, et al.
Veröffentlicht: (2025)
LeanConjecturer: Automatic Generation of Mathematical Conjectures for Theorem Proving
von: Onda, Naoto, et al.
Veröffentlicht: (2025)
von: Onda, Naoto, et al.
Veröffentlicht: (2025)
Lean Atlas: An Integrated Proof Environment for Scalable Human-AI Collaborative Formalization
von: Yanahama, Banri, et al.
Veröffentlicht: (2026)
von: Yanahama, Banri, et al.
Veröffentlicht: (2026)
SPECA: Specification-to-Checklist Agentic Auditing for Multi-Implementation Systems -- A Case Study on Ethereum Clients
von: Kamba, Masato, et al.
Veröffentlicht: (2026)
von: Kamba, Masato, et al.
Veröffentlicht: (2026)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
Deterministic Policy Gradient Primal-Dual Methods for Continuous-Space Constrained MDPs
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
von: Rozada, Sergio, et al.
Veröffentlicht: (2024)
Beyond Code Reasoning: Specification-Anchored Auditing of Multi-Implementation Distributed Protocols
von: Kamba, Masato, et al.
Veröffentlicht: (2026)
von: Kamba, Masato, et al.
Veröffentlicht: (2026)
Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
von: Hashimoto, Koji, et al.
Veröffentlicht: (2024)
von: Hashimoto, Koji, et al.
Veröffentlicht: (2024)
Decomposition of Equivariant Maps via Invariant Maps: Application to Universal Approximation under Symmetry
von: Sannai, Akiyoshi, et al.
Veröffentlicht: (2024)
von: Sannai, Akiyoshi, et al.
Veröffentlicht: (2024)
A unified Fourier slice method to derive ridgelet transform for a variety of depth-2 neural networks
von: Sonoda, Sho, et al.
Veröffentlicht: (2024)
von: Sonoda, Sho, et al.
Veröffentlicht: (2024)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2026)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2026)
MK2 at PBIG Competition: A Prompt Generation Solution
von: Xu, Yuzheng, et al.
Veröffentlicht: (2025)
von: Xu, Yuzheng, et al.
Veröffentlicht: (2025)
Prover Agent: An Agent-Based Framework for Formal Mathematical Proofs
von: Baba, Kaito, et al.
Veröffentlicht: (2025)
von: Baba, Kaito, et al.
Veröffentlicht: (2025)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
von: Nakao, Tomoki, et al.
Veröffentlicht: (2025)
von: Nakao, Tomoki, et al.
Veröffentlicht: (2025)
Self Iterative Label Refinement via Robust Unlabeled Learning
von: Asano, Hikaru, et al.
Veröffentlicht: (2025)
von: Asano, Hikaru, et al.
Veröffentlicht: (2025)
Multi-Agent Behavior Retrieval: Retrieval-Augmented Policy Training for Cooperative Push Manipulation by Mobile Robots
von: Kuroki, So, et al.
Veröffentlicht: (2023)
von: Kuroki, So, et al.
Veröffentlicht: (2023)
LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
von: Kuroki, So, et al.
Veröffentlicht: (2025)
von: Kuroki, So, et al.
Veröffentlicht: (2025)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
von: Bai, Qinbo, et al.
Veröffentlicht: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Why High-rank Neural Networks Generalize?: An Algebraic Framework with RKHSs
von: Hashimoto, Yuka, et al.
Veröffentlicht: (2025)
von: Hashimoto, Yuka, et al.
Veröffentlicht: (2025)
Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning
von: Sonoda, Sho, et al.
Veröffentlicht: (2025)
von: Sonoda, Sho, et al.
Veröffentlicht: (2025)
Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant Machines
von: Sonoda, Sho, et al.
Veröffentlicht: (2024)
von: Sonoda, Sho, et al.
Veröffentlicht: (2024)
Sampled-Data Primal-Dual Gradient Dynamics in Model Predictive Control
von: Moriyasu, Ryuta, et al.
Veröffentlicht: (2024)
von: Moriyasu, Ryuta, et al.
Veröffentlicht: (2024)
Sequential Audit Sampling with Statistical Guarantees
von: Kato, Masahiro, et al.
Veröffentlicht: (2026)
von: Kato, Masahiro, et al.
Veröffentlicht: (2026)
A Batch Sequential Halving Algorithm without Performance Degradation
von: Koyamada, Sotetsu, et al.
Veröffentlicht: (2024)
von: Koyamada, Sotetsu, et al.
Veröffentlicht: (2024)
Theoretical Guarantees for Minimum Bayes Risk Decoding
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
Generalization Error Bounds for Picard-Type Operator Learning in Nonlinear Parabolic PDEs
von: Taniguchi, Koichi, et al.
Veröffentlicht: (2026)
von: Taniguchi, Koichi, et al.
Veröffentlicht: (2026)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
von: Asano, Hikaru, et al.
Veröffentlicht: (2026)
von: Asano, Hikaru, et al.
Veröffentlicht: (2026)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
von: Xu, Yuzheng, et al.
Veröffentlicht: (2026)
von: Xu, Yuzheng, et al.
Veröffentlicht: (2026)
When to Replan? An Adaptive Replanning Strategy for Autonomous Navigation using Deep Reinforcement Learning
von: Honda, Kohei, et al.
Veröffentlicht: (2023)
von: Honda, Kohei, et al.
Veröffentlicht: (2023)
Cross-lingual Transfer or Machine Translation? On Data Augmentation for Monolingual Semantic Textual Similarity
von: Hoshino, Sho, et al.
Veröffentlicht: (2024)
von: Hoshino, Sho, et al.
Veröffentlicht: (2024)
Two-bridge links and stable maps into the plane
von: Ichihara, Kazuhiro, et al.
Veröffentlicht: (2024)
von: Ichihara, Kazuhiro, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2025) -
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2026) -
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024) -
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
von: Kasaura, Kazumi, et al.
Veröffentlicht: (2025) -
LeanConjecturer: Automatic Generation of Mathematical Conjectures for Theorem Proving
von: Onda, Naoto, et al.
Veröffentlicht: (2025)