Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
Fuente:
arXiv
Guardado en:
| Autores principales: | Kitamura, Toshinori, Kozuno, Tadashi, Kumagai, Wataru, Hoshino, Kenta, Hosoe, Yohei, Kasaura, Kazumi, Hamaya, Masashi, Parmas, Paavo, Matsuo, Yutaka |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
por: Kitamura, Toshinori, et al.
Publicado: (2025)
por: Kitamura, Toshinori, et al.
Publicado: (2025)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
por: Nishimori, Soichiro, et al.
Publicado: (2026)
por: Nishimori, Soichiro, et al.
Publicado: (2026)
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
por: Nakao, Tomoki, et al.
Publicado: (2025)
por: Nakao, Tomoki, et al.
Publicado: (2025)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
por: Kitamura, Toshinori, et al.
Publicado: (2024)
por: Kitamura, Toshinori, et al.
Publicado: (2024)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
por: Onoda, Ku, et al.
Publicado: (2026)
por: Onoda, Ku, et al.
Publicado: (2026)
Symmetry-aware Reinforcement Learning for Robotic Assembly under Partial Observability with a Soft Wrist
por: Nguyen, Hai, et al.
Publicado: (2024)
por: Nguyen, Hai, et al.
Publicado: (2024)
Compactness in Constructive Mathematics via Affine Logic
por: Kasaura, Kazumi
Publicado: (2026)
por: Kasaura, Kazumi
Publicado: (2026)
Homotopy-Aware Multi-Agent Path Planning on Plane
por: Kasaura, Kazumi
Publicado: (2023)
por: Kasaura, Kazumi
Publicado: (2023)
Generation of Geodesics with Actor-Critic Reinforcement Learning to Predict Midpoints
por: Kasaura, Kazumi
Publicado: (2024)
por: Kasaura, Kazumi
Publicado: (2024)
Double Horizon Model-Based Policy Optimization
por: Kubo, Akihiro, et al.
Publicado: (2025)
por: Kubo, Akihiro, et al.
Publicado: (2025)
Finite-Time Regret Analysis of Retry-Aware Bandits
por: Tong, Bingkui, et al.
Publicado: (2026)
por: Tong, Bingkui, et al.
Publicado: (2026)
Self Iterative Label Refinement via Robust Unlabeled Learning
por: Asano, Hikaru, et al.
Publicado: (2025)
por: Asano, Hikaru, et al.
Publicado: (2025)
Multi-Agent Behavior Retrieval: Retrieval-Augmented Policy Training for Cooperative Push Manipulation by Mobile Robots
por: Kuroki, So, et al.
Publicado: (2023)
por: Kuroki, So, et al.
Publicado: (2023)
An Electromagnetism-Inspired Method for Estimating In-Grasp Torque from Visuotactile Sensors
por: Fuchioka, Yuni, et al.
Publicado: (2024)
por: Fuchioka, Yuni, et al.
Publicado: (2024)
Uniformly boundedness of finite Morse index solutions to semilinear elliptic equations with rapidly growing nonlinearities in two dimensions
por: Kumagai, Kenta
Publicado: (2025)
por: Kumagai, Kenta
Publicado: (2025)
Bifurcation diagrams for semilinear elliptic equations with singular weights in two dimensions
por: Kumagai, Kenta
Publicado: (2024)
por: Kumagai, Kenta
Publicado: (2024)
Bifurcation diagrams of semilinear elliptic equations for supercritical nonlinearities in two dimensions
por: Kumagai, Kenta
Publicado: (2024)
por: Kumagai, Kenta
Publicado: (2024)
Classification of bifurcation structure for semilinear elliptic equations in a ball
por: Kumagai, Kenta
Publicado: (2025)
por: Kumagai, Kenta
Publicado: (2025)
Assessing Human Intelligence Augmentation Strategies Using Brain Machine Interfaces and Brain Organoids in the Era of AI Advancement
por: Kitamura, Kenta
Publicado: (2025)
por: Kitamura, Kenta
Publicado: (2025)
Stochastic Safety-critical Control Compensating Safety Probability for Marine Vessel Tracking
por: Matsuo, Too, et al.
Publicado: (2026)
por: Matsuo, Too, et al.
Publicado: (2026)
Efficient Constrained Signal Reconstruction by Randomized Epigraphical Projection
por: Ono, Shunsuke
Publicado: (2018)
por: Ono, Shunsuke
Publicado: (2018)
Lean Formalization of Generalization Error Bound by Rademacher Complexity and Dudley's Entropy Integral
por: Sonoda, Sho, et al.
Publicado: (2025)
por: Sonoda, Sho, et al.
Publicado: (2025)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
por: Watanabe, Yusuke, et al.
Publicado: (2026)
por: Watanabe, Yusuke, et al.
Publicado: (2026)
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
por: Asano, Hikaru, et al.
Publicado: (2026)
por: Asano, Hikaru, et al.
Publicado: (2026)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
por: Xu, Yuzheng, et al.
Publicado: (2026)
por: Xu, Yuzheng, et al.
Publicado: (2026)
When to Replan? An Adaptive Replanning Strategy for Autonomous Navigation using Deep Reinforcement Learning
por: Honda, Kohei, et al.
Publicado: (2023)
por: Honda, Kohei, et al.
Publicado: (2023)
A note on quantum subgroups of free quantum groups
por: Hoshino, Mao, et al.
Publicado: (2024)
por: Hoshino, Mao, et al.
Publicado: (2024)
A Multi‐Objective Evolutionary Algorithm Based on Bilayered Decomposition for Constrained Multi‐Objective Optimization
por: Yusuke Yasuda, et al.
Publicado: (2024)
por: Yusuke Yasuda, et al.
Publicado: (2024)
Pose Estimation of a Cable-Driven Serpentine Manipulator Utilizing Intrinsic Dynamics via Physical Reservoir Computing
por: Tanaka, Kazutoshi, et al.
Publicado: (2025)
por: Tanaka, Kazutoshi, et al.
Publicado: (2025)
Comments on resolution of nonassociativity in SFT- an example from axioms of BCFT-
por: Matsuo Yutaka
Publicado: (2002)
por: Matsuo Yutaka
Publicado: (2002)
Singular solutions and bifurcation diagram of semilinear elliptic equations with general nonlinearity in two dimensions
por: Kikuchi, Hiroaki, et al.
Publicado: (2025)
por: Kikuchi, Hiroaki, et al.
Publicado: (2025)
Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph Form
por: Wang, Xuefeng, et al.
Publicado: (2026)
por: Wang, Xuefeng, et al.
Publicado: (2026)
Solving Multi-Agent Safe Optimal Control with Distributed Epigraph Form MARL
por: Zhang, Songyuan, et al.
Publicado: (2025)
por: Zhang, Songyuan, et al.
Publicado: (2025)
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
por: Kasaura, Kazumi, et al.
Publicado: (2025)
por: Kasaura, Kazumi, et al.
Publicado: (2025)
LeanConjecturer: Automatic Generation of Mathematical Conjectures for Theorem Proving
por: Onda, Naoto, et al.
Publicado: (2025)
por: Onda, Naoto, et al.
Publicado: (2025)
Greek ΜΝΗΣΘΗ and Aramaic DKYR in the Near East: A Comparative Epigraphic Study
por: Sebastien Mazurek
Publicado: (2025)
por: Sebastien Mazurek
Publicado: (2025)
Control Barrier Functions for Stochastic Systems and Safety-critical Control Designs
por: Nishimura, Yuki, et al.
Publicado: (2022)
por: Nishimura, Yuki, et al.
Publicado: (2022)
LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
por: Kuroki, So, et al.
Publicado: (2025)
por: Kuroki, So, et al.
Publicado: (2025)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
por: Fiegel, Come, et al.
Publicado: (2026)
por: Fiegel, Come, et al.
Publicado: (2026)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
por: Fiegel, Côme, et al.
Publicado: (2026)
por: Fiegel, Côme, et al.
Publicado: (2026)
Ejemplares similares
-
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
por: Kitamura, Toshinori, et al.
Publicado: (2025) -
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
por: Nishimori, Soichiro, et al.
Publicado: (2026) -
Symmetry-Breaking in Multi-Agent Navigation: Winding Number-Aware MPC with a Learned Topological Strategy
por: Nakao, Tomoki, et al.
Publicado: (2025) -
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
por: Kitamura, Toshinori, et al.
Publicado: (2024) -
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
por: Onoda, Ku, et al.
Publicado: (2026)