Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds
Fuente:
arXiv
Guardado en:
| Autores principales: | Pariente, Yaacov, Indelman, Vadim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Online Risk-Averse Planning in POMDPs Using Iterated CVaR Value Function
por: Pariente, Yaacov, et al.
Publicado: (2026)
por: Pariente, Yaacov, et al.
Publicado: (2026)
Simplification of Risk Averse POMDPs with Performance Guarantees
por: Pariente, Yaacov, et al.
Publicado: (2024)
por: Pariente, Yaacov, et al.
Publicado: (2024)
Bounding Conditional Value-at-Risk via Auxiliary Distributions with Bounded Discrepancies
por: Pariente, Yaacov, et al.
Publicado: (2025)
por: Pariente, Yaacov, et al.
Publicado: (2025)
POMDPPlanners: Open-Source Package for POMDP Planning
por: Pariente, Yaacov, et al.
Publicado: (2026)
por: Pariente, Yaacov, et al.
Publicado: (2026)
Towards Optimal Performance and Action Consistency Guarantees in Dec-POMDPs with Inconsistent Beliefs and Limited Communication
por: Shimron, Moshe Rafaeli, et al.
Publicado: (2025)
por: Shimron, Moshe Rafaeli, et al.
Publicado: (2025)
Online POMDP Planning with Anytime Deterministic Optimality Guarantees
por: Barenboim, Moran, et al.
Publicado: (2023)
por: Barenboim, Moran, et al.
Publicado: (2023)
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
por: Mead, Harry, et al.
Publicado: (2025)
por: Mead, Harry, et al.
Publicado: (2025)
A Martingale approach to continuous Portfolio Optimization under CVaR like constraints
por: Lelong, Jérôme, et al.
Publicado: (2025)
por: Lelong, Jérôme, et al.
Publicado: (2025)
Measurement Simplification in ρ-POMDP with Performance Guarantees
por: Yotam, Tom, et al.
Publicado: (2023)
por: Yotam, Tom, et al.
Publicado: (2023)
No Compromise in Solution Quality: Speeding Up Belief-dependent Continuous POMDPs via Adaptive Multilevel Simplification
por: Zhitnikov, Andrey, et al.
Publicado: (2023)
por: Zhitnikov, Andrey, et al.
Publicado: (2023)
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
por: Wang, Kevin, et al.
Publicado: (2026)
por: Wang, Kevin, et al.
Publicado: (2026)
Conditional Performance Guarantee for Large Reasoning Models
por: Huang, Jianguo, et al.
Publicado: (2026)
por: Huang, Jianguo, et al.
Publicado: (2026)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
por: Yan, Hedong
Publicado: (2025)
por: Yan, Hedong
Publicado: (2025)
Anytime Probabilistically Constrained Provably Convergent Online Belief Space Planning
por: Zhitnikov, Andrey, et al.
Publicado: (2024)
por: Zhitnikov, Andrey, et al.
Publicado: (2024)
A Measure-Theoretic Axiomatisation of Causality
por: Park, Junhyung, et al.
Publicado: (2023)
por: Park, Junhyung, et al.
Publicado: (2023)
On the Provable Performance Guarantee of Efficient Reasoning Models
por: Zeng, Hao, et al.
Publicado: (2025)
por: Zeng, Hao, et al.
Publicado: (2025)
Simplifying Complex Observation Models in Continuous POMDP Planning with Probabilistic Guarantees and Practice
por: Lev-Yehudi, Idan, et al.
Publicado: (2023)
por: Lev-Yehudi, Idan, et al.
Publicado: (2023)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
por: Muni, Aneri, et al.
Publicado: (2026)
por: Muni, Aneri, et al.
Publicado: (2026)
Adaptive Insurance Reserving with CVaR-Constrained Reinforcement Learning under Macroeconomic Regimes
por: Dong, Stella C.
Publicado: (2025)
por: Dong, Stella C.
Publicado: (2025)
The Geometry of Knowing: From Possibilistic Ignorance to Probabilistic Certainty -- A Measure-Theoretic Framework for Epistemic Convergence
por: Jah, Moriba Kemessia
Publicado: (2026)
por: Jah, Moriba Kemessia
Publicado: (2026)
Guaranteed Recovery of Unambiguous Clusters
por: Mazooji, Kayvon, et al.
Publicado: (2025)
por: Mazooji, Kayvon, et al.
Publicado: (2025)
Online Learning with Unknown Constraints
por: Sridharan, Karthik, et al.
Publicado: (2024)
por: Sridharan, Karthik, et al.
Publicado: (2024)
Conformal Policy Control
por: Prinster, Drew, et al.
Publicado: (2026)
por: Prinster, Drew, et al.
Publicado: (2026)
Distributionally Robust Safety Verification of Neural Networks via Worst-Case CVaR
por: Kishida, Masako
Publicado: (2025)
por: Kishida, Masako
Publicado: (2025)
Finite-Time Analysis of MCTS in Continuous POMDP Planning
por: Kong, Da, et al.
Publicado: (2026)
por: Kong, Da, et al.
Publicado: (2026)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
por: Boudart, Pierre, et al.
Publicado: (2026)
por: Boudart, Pierre, et al.
Publicado: (2026)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
por: Hellström, Fredrik, et al.
Publicado: (2023)
por: Hellström, Fredrik, et al.
Publicado: (2023)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
por: Zhang, Bohan, et al.
Publicado: (2025)
por: Zhang, Bohan, et al.
Publicado: (2025)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
por: Chen, Fan, et al.
Publicado: (2025)
por: Chen, Fan, et al.
Publicado: (2025)
From Thomas Bayes to Big Data: On the feasibility of being a subjective Bayesian
por: Ritov, Ya'acov
Publicado: (2025)
por: Ritov, Ya'acov
Publicado: (2025)
No need for an oracle: the nonparametric maximum likelihood decision in the compound decision problem is minimax
por: Ritov, Ya'acov
Publicado: (2023)
por: Ritov, Ya'acov
Publicado: (2023)
A mixture of a normal distribution with random mean and variance -- Examples of inconsistency of maximum likelihood estimates
por: Ritov, Ya'acov
Publicado: (2024)
por: Ritov, Ya'acov
Publicado: (2024)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
por: Zhang, Huiming, et al.
Publicado: (2026)
por: Zhang, Huiming, et al.
Publicado: (2026)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
por: Lattimore, Tor
Publicado: (2026)
por: Lattimore, Tor
Publicado: (2026)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
por: Baharav, Tavor Z., et al.
Publicado: (2025)
por: Baharav, Tavor Z., et al.
Publicado: (2025)
Risk Analysis and Design Against Adversarial Actions
por: Campi, Marco C., et al.
Publicado: (2025)
por: Campi, Marco C., et al.
Publicado: (2025)
Le Cam Distortion: A Decision-Theoretic Framework for Robust Transfer Learning
por: Akdemir, Deniz
Publicado: (2025)
por: Akdemir, Deniz
Publicado: (2025)
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
por: Hao, Sai, et al.
Publicado: (2026)
por: Hao, Sai, et al.
Publicado: (2026)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
por: Yu, Lijia, et al.
Publicado: (2025)
por: Yu, Lijia, et al.
Publicado: (2025)
Conformal Risk Control
por: Angelopoulos, Anastasios N., et al.
Publicado: (2022)
por: Angelopoulos, Anastasios N., et al.
Publicado: (2022)
Ejemplares similares
-
Online Risk-Averse Planning in POMDPs Using Iterated CVaR Value Function
por: Pariente, Yaacov, et al.
Publicado: (2026) -
Simplification of Risk Averse POMDPs with Performance Guarantees
por: Pariente, Yaacov, et al.
Publicado: (2024) -
Bounding Conditional Value-at-Risk via Auxiliary Distributions with Bounded Discrepancies
por: Pariente, Yaacov, et al.
Publicado: (2025) -
POMDPPlanners: Open-Source Package for POMDP Planning
por: Pariente, Yaacov, et al.
Publicado: (2026) -
Towards Optimal Performance and Action Consistency Guarantees in Dec-POMDPs with Inconsistent Beliefs and Limited Communication
por: Shimron, Moshe Rafaeli, et al.
Publicado: (2025)