Maximum Entropy Exploration Without the Rollouts
Fuente:
arXiv
Saved in:
| Main Authors: | Adamczyk, Jacob, Kamoski, Adam, Kulkarni, Rahul V. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploration Behavior of Untrained Policies
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
Thermodynamics of Reinforcement Learning Curricula
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
EVAL: EigenVector-based Average-reward Learning
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Boosting Soft Q-Learning by Bounding
by: Adamczyk, Jacob, et al.
Published: (2024)
by: Adamczyk, Jacob, et al.
Published: (2024)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Inferring Transition Dynamics from Value Functions
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
by: Li, Hongming, et al.
Published: (2024)
by: Li, Hongming, et al.
Published: (2024)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
by: Hu, Jiajun, et al.
Published: (2026)
by: Hu, Jiajun, et al.
Published: (2026)
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
by: Zeng, Anxiang, et al.
Published: (2025)
by: Zeng, Anxiang, et al.
Published: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
by: Yoo, Seungwoo, et al.
Published: (2026)
by: Yoo, Seungwoo, et al.
Published: (2026)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
by: Dong, Xiaoyi, et al.
Published: (2025)
by: Dong, Xiaoyi, et al.
Published: (2025)
Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
by: Ma, Jiyong
Published: (2025)
by: Ma, Jiyong
Published: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
by: Sanokowski, Sebastian, et al.
Published: (2025)
by: Sanokowski, Sebastian, et al.
Published: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
by: Bi, Jinhe, et al.
Published: (2026)
by: Bi, Jinhe, et al.
Published: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
Emergency Preemption Without Online Exploration: A Decision Transformer Approach
by: Su, Haoran, et al.
Published: (2026)
by: Su, Haoran, et al.
Published: (2026)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
by: Li, Zeqiao, et al.
Published: (2026)
by: Li, Zeqiao, et al.
Published: (2026)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
by: Yoon, Sangwoong, et al.
Published: (2024)
by: Yoon, Sangwoong, et al.
Published: (2024)
RTMC: Step-Level Credit Assignment via Rollout Trees
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
ARROW: An Adaptive Rollout and Routing Method for Global Weather Forecasting
by: Tian, Jindong, et al.
Published: (2025)
by: Tian, Jindong, et al.
Published: (2025)
ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Evaluating machine learning models for predicting pesticide toxicity to honey bees
by: Adamczyk, Jakub, et al.
Published: (2025)
by: Adamczyk, Jakub, et al.
Published: (2025)
Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
by: Praski, Mateusz, et al.
Published: (2025)
by: Praski, Mateusz, et al.
Published: (2025)
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
by: Chen, Xinzhu, et al.
Published: (2025)
by: Chen, Xinzhu, et al.
Published: (2025)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
Entropy-Aware Model Initialization for Effective Exploration in Deep Reinforcement Learning
by: Jang, Sooyoung, et al.
Published: (2021)
by: Jang, Sooyoung, et al.
Published: (2021)
EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
by: Block, Adam, et al.
Published: (2025)
by: Block, Adam, et al.
Published: (2025)
Learning to Embed Distributions via Maximum Kernel Entropy
by: Kachaiev, Oleksii, et al.
Published: (2024)
by: Kachaiev, Oleksii, et al.
Published: (2024)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes
by: Middelhuis, Jeroen, et al.
Published: (2025)
by: Middelhuis, Jeroen, et al.
Published: (2025)
Similar Items
-
Exploration Behavior of Untrained Policies
by: Adamczyk, Jacob
Published: (2025) -
Thermodynamics of Reinforcement Learning Curricula
by: Adamczyk, Jacob, et al.
Published: (2026) -
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025) -
EVAL: EigenVector-based Average-reward Learning
by: Adamczyk, Jacob, et al.
Published: (2025) -
Boosting Soft Q-Learning by Bounding
by: Adamczyk, Jacob, et al.
Published: (2024)