Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
Fuente:
arXiv
Saved in:
| Main Authors: | Bolland, Adrien, Lambrechts, Gaspard, Ernst, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maximum-Entropy Exploration with Future State-Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2026)
by: Bolland, Adrien, et al.
Published: (2026)
Informed POMDP: Leveraging Additional Information in Model-Based RL
by: Lambrechts, Gaspard, et al.
Published: (2023)
by: Lambrechts, Gaspard, et al.
Published: (2023)
Behind the Myth of Exploration in Policy Gradients
by: Bolland, Adrien, et al.
Published: (2024)
by: Bolland, Adrien, et al.
Published: (2024)
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
by: Lambrechts, Gaspard, et al.
Published: (2025)
by: Lambrechts, Gaspard, et al.
Published: (2025)
Parallelizing Autoregressive Generation with Variational State Space Models
by: Lambrechts, Gaspard, et al.
Published: (2024)
by: Lambrechts, Gaspard, et al.
Published: (2024)
Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
by: Ebi, Daniel, et al.
Published: (2025)
by: Ebi, Daniel, et al.
Published: (2025)
Cost Estimation in Unit Commitment Problems Using Simulation-Based Inference
by: Pirlet, Matthias, et al.
Published: (2024)
by: Pirlet, Matthias, et al.
Published: (2024)
Gym-TORAX: Open-source software for integrating reinforcement learning with plasma control simulators in tokamak research
by: Mouchamps, Antoine, et al.
Published: (2025)
by: Mouchamps, Antoine, et al.
Published: (2025)
Parallelizable memory recurrent units
by: De Geeter, Florent, et al.
Published: (2026)
by: De Geeter, Florent, et al.
Published: (2026)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Reinforcement Learning for Efficient Design and Control Co-optimisation of Energy Systems
by: Cauz, Marine, et al.
Published: (2024)
by: Cauz, Marine, et al.
Published: (2024)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
LLMs Can Learn to Reason Via Off-Policy RL
by: Ritter, Daniel, et al.
Published: (2026)
by: Ritter, Daniel, et al.
Published: (2026)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Action-Free Offline-to-Online RL via Discretised State Policies
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
by: Neggatu, Natinael Solomon, et al.
Published: (2026)
Policy Learning for Off-Dynamics RL with Deficient Support
by: Van, Linh Le Pham, et al.
Published: (2024)
by: Van, Linh Le Pham, et al.
Published: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Efficient Off-Policy Learning for High-Dimensional Action Spaces
by: Otto, Fabian, et al.
Published: (2024)
by: Otto, Fabian, et al.
Published: (2024)
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data
by: Romeo, Carlo, et al.
Published: (2026)
by: Romeo, Carlo, et al.
Published: (2026)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
by: Admoni, Sahar, et al.
Published: (2025)
by: Admoni, Sahar, et al.
Published: (2025)
Optimal Control of Renewable Energy Communities subject to Network Peak Fees with Model Predictive Control and Reinforcement Learning Algorithms
by: Aittahar, Samy, et al.
Published: (2024)
by: Aittahar, Samy, et al.
Published: (2024)
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
by: Luo, Fan-Ming, et al.
Published: (2024)
by: Luo, Fan-Ming, et al.
Published: (2024)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL
by: Anahtarci, Berkay, et al.
Published: (2024)
by: Anahtarci, Berkay, et al.
Published: (2024)
A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and Preconditioning
by: Wu, Zechen, et al.
Published: (2025)
by: Wu, Zechen, et al.
Published: (2025)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Similar Items
-
Maximum-Entropy Exploration with Future State-Action Visitation Measures
by: Bolland, Adrien, et al.
Published: (2026) -
Informed POMDP: Leveraging Additional Information in Model-Based RL
by: Lambrechts, Gaspard, et al.
Published: (2023) -
Behind the Myth of Exploration in Policy Gradients
by: Bolland, Adrien, et al.
Published: (2024) -
A Theoretical Justification for Asymmetric Actor-Critic Algorithms
by: Lambrechts, Gaspard, et al.
Published: (2025) -
Parallelizing Autoregressive Generation with Variational State Space Models
by: Lambrechts, Gaspard, et al.
Published: (2024)