Data-Efficient Safe Policy Improvement Using Parametric Structure
Fuente:
arXiv
Guardado en:
| Autores principales: | Engelen, Kasper, Pérez, Guillermo A., Suilen, Marnix |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Analyzing Value Functions of States in Parametric Markov Chains
por: Engelen, Kasper, et al.
Publicado: (2025)
por: Engelen, Kasper, et al.
Publicado: (2025)
On the Complexity of Robust Markov Decision Processes and Bisimulation Metrics
por: Suilen, Marnix, et al.
Publicado: (2026)
por: Suilen, Marnix, et al.
Publicado: (2026)
Imprecise Probabilities Meet Partial Observability: Game Semantics for Robust POMDPs
por: Bovy, Eline M., et al.
Publicado: (2024)
por: Bovy, Eline M., et al.
Publicado: (2024)
Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability
por: Bovy, Eline M., et al.
Publicado: (2025)
por: Bovy, Eline M., et al.
Publicado: (2025)
Robust Markov Decision Processes: A Place Where AI and Formal Methods Meet
por: Suilen, Marnix, et al.
Publicado: (2024)
por: Suilen, Marnix, et al.
Publicado: (2024)
Institutional AI Sovereignty Through Gateway Architecture: Implementation Report from Fontys ICT
por: Huijts, Ruud, et al.
Publicado: (2025)
por: Huijts, Ruud, et al.
Publicado: (2025)
Vibe coding before the trend
por: van Bokhorst, Leon, et al.
Publicado: (2026)
por: van Bokhorst, Leon, et al.
Publicado: (2026)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
por: Galesloot, Maris F. L., et al.
Publicado: (2024)
por: Galesloot, Maris F. L., et al.
Publicado: (2024)
Deep SPI: Safe Policy Improvement via World Models
por: Delgrange, Florent, et al.
Publicado: (2025)
por: Delgrange, Florent, et al.
Publicado: (2025)
Policy-Aware Generative AI for Safe, Auditable Data Access Governance
por: Mandalawi, Shames Al, et al.
Publicado: (2025)
por: Mandalawi, Shames Al, et al.
Publicado: (2025)
Safe Explicable Policy Search
por: Hanni, Akkamahadevi, et al.
Publicado: (2025)
por: Hanni, Akkamahadevi, et al.
Publicado: (2025)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
por: Li, Hao, et al.
Publicado: (2026)
por: Li, Hao, et al.
Publicado: (2026)
Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
por: Kiyohara, Haruka, et al.
Publicado: (2025)
por: Kiyohara, Haruka, et al.
Publicado: (2025)
Policy-Space Search: Equivalences, Improvements, and Compression
por: Messa, Frederico, et al.
Publicado: (2024)
por: Messa, Frederico, et al.
Publicado: (2024)
SafeAR: Safe Algorithmic Recourse by Risk-Aware Policies
por: Wu, Haochen, et al.
Publicado: (2023)
por: Wu, Haochen, et al.
Publicado: (2023)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
por: Jin, Yang, et al.
Publicado: (2025)
por: Jin, Yang, et al.
Publicado: (2025)
Safe Equilibrium Policy Optimization for Strategic Agent Policies
por: Arumugam, Karthika, et al.
Publicado: (2026)
por: Arumugam, Karthika, et al.
Publicado: (2026)
Offline Safe Policy Optimization From Heterogeneous Feedback
por: Gong, Ze, et al.
Publicado: (2025)
por: Gong, Ze, et al.
Publicado: (2025)
Learning State-Dependent Policy Parametrizations for Dynamic Technician Routing with Rework
por: Stein, Jonas, et al.
Publicado: (2024)
por: Stein, Jonas, et al.
Publicado: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
por: Anisimov, Maksim, et al.
Publicado: (2026)
por: Anisimov, Maksim, et al.
Publicado: (2026)
Safe Deep Policy Adaptation
por: Xiao, Wenli, et al.
Publicado: (2023)
por: Xiao, Wenli, et al.
Publicado: (2023)
Anytime Safe PAC Efficient Reasoning
por: Yu, Chengyao, et al.
Publicado: (2026)
por: Yu, Chengyao, et al.
Publicado: (2026)
Data-Efficient On-Policy Distillation for Automatic Speech Recognition
por: Lin, Yu, et al.
Publicado: (2026)
por: Lin, Yu, et al.
Publicado: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
por: Li, Xiang, et al.
Publicado: (2026)
por: Li, Xiang, et al.
Publicado: (2026)
Safe Planning and Policy Optimization via World Model Learning
por: Latyshev, Artem, et al.
Publicado: (2025)
por: Latyshev, Artem, et al.
Publicado: (2025)
Composing Reinforcement Learning Policies, with Formal Guarantees
por: Delgrange, Florent, et al.
Publicado: (2024)
por: Delgrange, Florent, et al.
Publicado: (2024)
Safe Exploration via Policy Priors
por: Wendl, Manuel, et al.
Publicado: (2026)
por: Wendl, Manuel, et al.
Publicado: (2026)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
por: Burnwal, Returaj, et al.
Publicado: (2025)
por: Burnwal, Returaj, et al.
Publicado: (2025)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
por: Xue, Ruiqi, et al.
Publicado: (2026)
por: Xue, Ruiqi, et al.
Publicado: (2026)
STRIVE: Structured Reasoning for Self-Improvement in Claim Verification
por: Gong, Haisong, et al.
Publicado: (2025)
por: Gong, Haisong, et al.
Publicado: (2025)
Provable and Practical In-Context Policy Optimization for Self-Improvement
por: Yu, Tianrun, et al.
Publicado: (2026)
por: Yu, Tianrun, et al.
Publicado: (2026)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
por: Liu, Xuefeng, et al.
Publicado: (2023)
por: Liu, Xuefeng, et al.
Publicado: (2023)
Certifiably Robust Policies for Uncertain Parametric Environments
por: Schnitzer, Yannik, et al.
Publicado: (2024)
por: Schnitzer, Yannik, et al.
Publicado: (2024)
A PSPACE Algorithm for Almost-Sure Rabin Objectives in Multi-Environment MDPs
por: Suilen, Marnix, et al.
Publicado: (2024)
por: Suilen, Marnix, et al.
Publicado: (2024)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
por: Xu, Peiyang, et al.
Publicado: (2025)
por: Xu, Peiyang, et al.
Publicado: (2025)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
por: Tayal, Mumuksh, et al.
Publicado: (2026)
por: Tayal, Mumuksh, et al.
Publicado: (2026)
Group and Shuffle: Efficient Structured Orthogonal Parametrization
por: Gorbunov, Mikhail, et al.
Publicado: (2024)
por: Gorbunov, Mikhail, et al.
Publicado: (2024)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
por: Lee, Chi-Chang, et al.
Publicado: (2025)
por: Lee, Chi-Chang, et al.
Publicado: (2025)
EasyInsert: A Data-Efficient and Generalizable Insertion Policy
por: Li, Guanghe, et al.
Publicado: (2025)
por: Li, Guanghe, et al.
Publicado: (2025)
Active Policy Improvement from Multiple Black-box Oracles
por: Liu, Xuefeng, et al.
Publicado: (2023)
por: Liu, Xuefeng, et al.
Publicado: (2023)
Ejemplares similares
-
Analyzing Value Functions of States in Parametric Markov Chains
por: Engelen, Kasper, et al.
Publicado: (2025) -
On the Complexity of Robust Markov Decision Processes and Bisimulation Metrics
por: Suilen, Marnix, et al.
Publicado: (2026) -
Imprecise Probabilities Meet Partial Observability: Game Semantics for Robust POMDPs
por: Bovy, Eline M., et al.
Publicado: (2024) -
Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability
por: Bovy, Eline M., et al.
Publicado: (2025) -
Robust Markov Decision Processes: A Place Where AI and Formal Methods Meet
por: Suilen, Marnix, et al.
Publicado: (2024)