Entropy-regularized Point-based Value Iteration
Fuente:
arXiv
Saved in:
| Main Authors: | Delecki, Harrison, Vazquez-Chanlatte, Marcell, Yel, Esen, Wray, Kyle, Arnon, Tomer, Witwicki, Stefan, Kochenderfer, Mykel J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion-Based Failure Sampling for Evaluating Safety-Critical Autonomous Systems
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Diffusion Models for Safety Validation of Autonomous Driving Systems
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
$L^*LM$: Learning Automata from Examples using Natural Language Oracles
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
Semi-Markovian Planning to Coordinate Aerial and Maritime Medical Evacuation Platforms
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
A Semi-Decentralized Approach to Multiagent Control
by: Al-Husseini, Mahdi, et al.
Published: (2026)
by: Al-Husseini, Mahdi, et al.
Published: (2026)
Enhanced Importance Sampling through Latent Space Exploration in Normalizing Flows
by: Kruse, Liam A., et al.
Published: (2025)
by: Kruse, Liam A., et al.
Published: (2025)
Optimizing Task Completion Time Updates Using POMDPs
by: Eddy, Duncan, et al.
Published: (2026)
by: Eddy, Duncan, et al.
Published: (2026)
Constrained Hierarchical Monte Carlo Belief-State Planning
by: Jamgochian, Arec, et al.
Published: (2023)
by: Jamgochian, Arec, et al.
Published: (2023)
Predicting Future Spatiotemporal Occupancy Grids with Semantics for Autonomous Driving
by: Toyungyernsub, Maneekwan, et al.
Published: (2023)
by: Toyungyernsub, Maneekwan, et al.
Published: (2023)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025)
by: Chaubard, Francois, et al.
Published: (2025)
Risk-aware Meta-level Decision Making for Exploration Under Uncertainty
by: Ott, Joshua, et al.
Published: (2022)
by: Ott, Joshua, et al.
Published: (2022)
Large-Scale Multi-Robot Assembly Planning for Autonomous Manufacturing
by: Brown, Kyle, et al.
Published: (2023)
by: Brown, Kyle, et al.
Published: (2023)
Conditional Deep Generative Models for Belief State Planning
by: Bigeard, Antoine, et al.
Published: (2025)
by: Bigeard, Antoine, et al.
Published: (2025)
Using Language and Road Manuals to Inform Map Reconstruction for Autonomous Driving
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
Learning Formal Specifications from Membership and Preference Queries
by: Shah, Ameesh, et al.
Published: (2023)
by: Shah, Ameesh, et al.
Published: (2023)
Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments
by: Lange, Bernard, et al.
Published: (2023)
by: Lange, Bernard, et al.
Published: (2023)
Watercraft as Overwater Ambulance Exchange Points to Enhance Aeromedical Evacuation
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)
by: Chaubard, Francois, et al.
Published: (2024)
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
Robust Planning for Autonomous Vehicles with Diffusion-Based Failure Samplers
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
Addressing Myopic Constrained POMDP Planning with Recursive Dual Ascent
by: Stocco, Paula, et al.
Published: (2024)
by: Stocco, Paula, et al.
Published: (2024)
Hierarchical Framework for Optimizing Wildfire Surveillance and Suppression using Human-Autonomous Teaming
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
by: Eddy, Duncan, et al.
Published: (2024)
by: Eddy, Duncan, et al.
Published: (2024)
BetaZero: Belief-State Planning for Long-Horizon POMDPs using Learned Approximations
by: Moss, Robert J., et al.
Published: (2023)
by: Moss, Robert J., et al.
Published: (2023)
Improving the Resilience of Quadrotors in Underground Environments by Combining Learning-based and Safety Controllers
by: Ward, Isaac Ronald, et al.
Published: (2025)
by: Ward, Isaac Ronald, et al.
Published: (2025)
On Technique Identification and Threat-Actor Attribution using LLMs and Embedding Models
by: Guru, Kyla, et al.
Published: (2025)
by: Guru, Kyla, et al.
Published: (2025)
A Taxonomy and Review of Algorithms for Modeling and Predicting Human Driver Behavior
by: Bhattacharyya, Raunak P., et al.
Published: (2020)
by: Bhattacharyya, Raunak P., et al.
Published: (2020)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
by: Lamparth, Max, et al.
Published: (2026)
by: Lamparth, Max, et al.
Published: (2026)
A New Strategy for Verifying Reach-Avoid Specifications in Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
by: Yalcinkaya, Beyazit, et al.
Published: (2025)
Graph Q-Learning for Combinatorial Optimization
by: Dax, Victoria M., et al.
Published: (2024)
by: Dax, Victoria M., et al.
Published: (2024)
ConstrainedZero: Chance-Constrained POMDP Planning using Learned Probabilistic Failure Surrogates and Adaptive Safety Constraints
by: Moss, Robert J., et al.
Published: (2024)
by: Moss, Robert J., et al.
Published: (2024)
Zono-Conformal Prediction: Zonotope-Based Uncertainty Quantification for Regression and Classification Tasks
by: Lützow, Laura, et al.
Published: (2025)
by: Lützow, Laura, et al.
Published: (2025)
Importance Sampling-Guided Meta-Training for Intelligent Agents in Highly Interactive Environments
by: Arief, Mansur, et al.
Published: (2024)
by: Arief, Mansur, et al.
Published: (2024)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
Semi‐Markovian planning to coordinate aerial and maritime medical evacuation platforms
by: Mahdi Al‐Husseini, et al.
Published: (2025)
by: Mahdi Al‐Husseini, et al.
Published: (2025)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
by: Fein, Daniel, et al.
Published: (2026)
by: Fein, Daniel, et al.
Published: (2026)
The FABRIC Strategy for Verifying Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
Similar Items
-
Diffusion-Based Failure Sampling for Evaluating Safety-Critical Autonomous Systems
by: Delecki, Harrison, et al.
Published: (2024) -
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024) -
Diffusion Models for Safety Validation of Autonomous Driving Systems
by: Wang, Juanran, et al.
Published: (2025) -
$L^*LM$: Learning Automata from Examples using Natural Language Oracles
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024) -
Semi-Markovian Planning to Coordinate Aerial and Maritime Medical Evacuation Platforms
by: Al-Husseini, Mahdi, et al.
Published: (2024)