Informed POMDP: Leveraging Additional Information in Model-Based RL
Fuente:
arXiv
Saved in:
| Main Authors: | Lambrechts, Gaspard, Bolland, Adrien, Ernst, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2024)
by: Bolland, Adrien, et al.
Published: (2024)
Behind the Myth of Exploration in Policy Gradients
by: Bolland, Adrien, et al.
Published: (2024)
by: Bolland, Adrien, et al.
Published: (2024)
Maximum-Entropy Exploration with Future State-Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2026)
by: Bolland, Adrien, et al.
Published: (2026)
Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
by: Ebi, Daniel, et al.
Published: (2025)
by: Ebi, Daniel, et al.
Published: (2025)
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
by: Lambrechts, Gaspard, et al.
Published: (2025)
by: Lambrechts, Gaspard, et al.
Published: (2025)
Parallelizing Autoregressive Generation with Variational State Space Models
by: Lambrechts, Gaspard, et al.
Published: (2024)
by: Lambrechts, Gaspard, et al.
Published: (2024)
Cost Estimation in Unit Commitment Problems Using Simulation-Based Inference
by: Pirlet, Matthias, et al.
Published: (2024)
by: Pirlet, Matthias, et al.
Published: (2024)
Gym-TORAX: Open-source software for integrating reinforcement learning with plasma control simulators in tokamak research
by: Mouchamps, Antoine, et al.
Published: (2025)
by: Mouchamps, Antoine, et al.
Published: (2025)
Parallelizable memory recurrent units
by: De Geeter, Florent, et al.
Published: (2026)
by: De Geeter, Florent, et al.
Published: (2026)
Reinforcement Learning for Efficient Design and Control Co-optimisation of Energy Systems
by: Cauz, Marine, et al.
Published: (2024)
by: Cauz, Marine, et al.
Published: (2024)
SOAP-RL: Sequential Option Advantage Propagation for Reinforcement Learning in POMDP Environments
by: Ishida, Shu, et al.
Published: (2024)
by: Ishida, Shu, et al.
Published: (2024)
Optimal Control of Renewable Energy Communities subject to Network Peak Fees with Model Predictive Control and Reinforcement Learning Algorithms
by: Aittahar, Samy, et al.
Published: (2024)
by: Aittahar, Samy, et al.
Published: (2024)
Discovery of Sustainable Refrigerants through Physics-Informed RL Fine-Tuning of Sequence Models
by: Goldszal, Adrien, et al.
Published: (2025)
by: Goldszal, Adrien, et al.
Published: (2025)
Deep Belief Markov Models for POMDP Inference
by: Arcieri, Giacomo, et al.
Published: (2025)
by: Arcieri, Giacomo, et al.
Published: (2025)
Learning POMDP World Models from Observations with Language-Model Priors
by: Six, Valentin, et al.
Published: (2026)
by: Six, Valentin, et al.
Published: (2026)
C-IDS: Solving Contextual POMDP via Information-Directed Objective
by: Shi, Chongyang, et al.
Published: (2026)
by: Shi, Chongyang, et al.
Published: (2026)
Rao-Blackwellized POMDP Planning
by: Lee, Jiho, et al.
Published: (2024)
by: Lee, Jiho, et al.
Published: (2024)
Reinforcement Learning to improve delta robot throws for sorting scrap metal
by: Louette, Arthur, et al.
Published: (2024)
by: Louette, Arthur, et al.
Published: (2024)
Expert-Guided POMDP Learning for Data-Efficient Modeling in Healthcare
by: Locatelli, Marco, et al.
Published: (2025)
by: Locatelli, Marco, et al.
Published: (2025)
Reinforcement Learning in POMDP's via Direct Gradient Ascent
by: Baxter, Jonathan, et al.
Published: (2025)
by: Baxter, Jonathan, et al.
Published: (2025)
Probing Dec-POMDP Reasoning in Cooperative MARL
by: Tessera, Kale-ab, et al.
Published: (2026)
by: Tessera, Kale-ab, et al.
Published: (2026)
Real-World Reinforcement Learning of Active Perception Behaviors
by: Hu, Edward S., et al.
Published: (2025)
by: Hu, Edward S., et al.
Published: (2025)
Learning Optimal Defender Strategies for CAGE-2 using a POMDP Model
by: Le, Duc Huy, et al.
Published: (2025)
by: Le, Duc Huy, et al.
Published: (2025)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
Anytime Incremental $ρ$POMDP Planning in Continuous Spaces
by: Benchetrit, Ron, et al.
Published: (2025)
by: Benchetrit, Ron, et al.
Published: (2025)
Meta-RL Induces Exploration in Language Agents
by: Jiang, Yulun, et al.
Published: (2025)
by: Jiang, Yulun, et al.
Published: (2025)
SOMBRL: Scalable and Optimistic Model-Based RL
by: Sukhija, Bhavya, et al.
Published: (2025)
by: Sukhija, Bhavya, et al.
Published: (2025)
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023)
by: Kraske, Benjamin D., et al.
Published: (2023)
Spike-based computation using classical recurrent neural networks
by: De Geeter, Florent, et al.
Published: (2023)
by: De Geeter, Florent, et al.
Published: (2023)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
by: Kausik, Chinmaya, et al.
Published: (2026)
by: Kausik, Chinmaya, et al.
Published: (2026)
Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
by: Feng, Yunzhen, et al.
Published: (2025)
by: Feng, Yunzhen, et al.
Published: (2025)
Federated Client Selection under Partial Visibility: A POMDP Approach with Spatio-Temporal Attention
by: Hou, Qijun, et al.
Published: (2026)
by: Hou, Qijun, et al.
Published: (2026)
How to craft a deep reinforcement learning policy for wind farm flow control
by: Kadoche, Elie, et al.
Published: (2025)
by: Kadoche, Elie, et al.
Published: (2025)
Acceleration Methods
by: d'Aspremont, Alexandre, et al.
Published: (2021)
by: d'Aspremont, Alexandre, et al.
Published: (2021)
Learning Explainable and Better Performing Representations of POMDP Strategies
by: Bork, Alexander, et al.
Published: (2024)
by: Bork, Alexander, et al.
Published: (2024)
Investigating simple target-covariate relationships for Chronos-2 and TabPFN-TS
by: Berthelier, Gaspard, et al.
Published: (2026)
by: Berthelier, Gaspard, et al.
Published: (2026)
Hierarchical Object-Oriented POMDP Planning for Object Rearrangement
by: Mangannavar, Rajesh, et al.
Published: (2024)
by: Mangannavar, Rajesh, et al.
Published: (2024)
From Ambiguity to Action: A POMDP Perspective on Partial Multi-Label Ambiguity and Its Horizon-One Resolution
by: Pan, Hanlin, et al.
Published: (2026)
by: Pan, Hanlin, et al.
Published: (2026)
Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL
by: Wagenmaker, Andrew, et al.
Published: (2024)
by: Wagenmaker, Andrew, et al.
Published: (2024)
Fair Reinforcement Learning Algorithm for PV Active Control in LV Distribution Networks
by: Vassallo, Maurizio, et al.
Published: (2024)
by: Vassallo, Maurizio, et al.
Published: (2024)
Similar Items
-
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2024) -
Behind the Myth of Exploration in Policy Gradients
by: Bolland, Adrien, et al.
Published: (2024) -
Maximum-Entropy Exploration with Future State-Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2026) -
Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
by: Ebi, Daniel, et al.
Published: (2025) -
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
by: Lambrechts, Gaspard, et al.
Published: (2025)