Exploration Behavior of Untrained Policies
Fuente:
arXiv
Saved in:
| Main Author: | Adamczyk, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
Inferring Transition Dynamics from Value Functions
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
Link Prediction with Untrained Message Passing Layers
by: Qarkaxhija, Lisi, et al.
Published: (2024)
by: Qarkaxhija, Lisi, et al.
Published: (2024)
Thermodynamics of Reinforcement Learning Curricula
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
RandomNet: Clustering Time Series Using Untrained Deep Neural Networks
by: Li, Xiaosheng, et al.
Published: (2024)
by: Li, Xiaosheng, et al.
Published: (2024)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
EVAL: EigenVector-based Average-reward Learning
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Boosting Soft Q-Learning by Bounding
by: Adamczyk, Jacob, et al.
Published: (2024)
by: Adamczyk, Jacob, et al.
Published: (2024)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control
by: Islam, SM Mazharul, et al.
Published: (2025)
by: Islam, SM Mazharul, et al.
Published: (2025)
Evaluating machine learning models for predicting pesticide toxicity to honey bees
by: Adamczyk, Jakub, et al.
Published: (2025)
by: Adamczyk, Jakub, et al.
Published: (2025)
Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
by: Praski, Mateusz, et al.
Published: (2025)
by: Praski, Mateusz, et al.
Published: (2025)
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
by: Wan, Zhenglin, et al.
Published: (2024)
by: Wan, Zhenglin, et al.
Published: (2024)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
How Log-Barrier Helps Exploration in Policy Optimization
by: Cesani, Leonardo, et al.
Published: (2026)
by: Cesani, Leonardo, et al.
Published: (2026)
Safe Exploration via Policy Priors
by: Wendl, Manuel, et al.
Published: (2026)
by: Wendl, Manuel, et al.
Published: (2026)
Symmetric Behavior Regularized Policy Optimization
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
by: Huang, Zhuoxu, et al.
Published: (2026)
by: Huang, Zhuoxu, et al.
Published: (2026)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Efficient On-Policy Reinforcement Learning via Exploration of Sparse Parameter Space
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
A Causal Lens for Learning Long-term Fair Policies
by: Lear, Jacob, et al.
Published: (2025)
by: Lear, Jacob, et al.
Published: (2025)
Learning Policy Representations for Steerable Behavior Synthesis
by: Li, Beiming, et al.
Published: (2026)
by: Li, Beiming, et al.
Published: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Policy Optimization for Personalized Interventions in Behavioral Health
by: Baek, Jackie, et al.
Published: (2023)
by: Baek, Jackie, et al.
Published: (2023)
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
by: Dunefsky, Jacob, et al.
Published: (2025)
by: Dunefsky, Jacob, et al.
Published: (2025)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
by: Mohamed, Faisal, et al.
Published: (2026)
by: Mohamed, Faisal, et al.
Published: (2026)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
Scheduled Curiosity-Deep Dyna-Q: Efficient Exploration for Dialog Policy Learning
by: Niu, Xuecheng, et al.
Published: (2024)
by: Niu, Xuecheng, et al.
Published: (2024)
Pragmatic Policy Development via Interpretable Behavior Cloning
by: Matsson, Anton, et al.
Published: (2025)
by: Matsson, Anton, et al.
Published: (2025)
Deep Causal Behavioral Policy Learning: Applications to Healthcare
by: Knecht, Jonas, et al.
Published: (2025)
by: Knecht, Jonas, et al.
Published: (2025)
From Parameters to Behaviors: Unsupervised Compression of the Policy Space
by: Tenedini, Davide, et al.
Published: (2025)
by: Tenedini, Davide, et al.
Published: (2025)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
by: Hu, Jiajun, et al.
Published: (2026)
by: Hu, Jiajun, et al.
Published: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
by: Bai, Qinxun, et al.
Published: (2025)
by: Bai, Qinxun, et al.
Published: (2025)
An Optimal Policy for Learning Controllable Dynamics by Exploration
by: Loxley, Peter N.
Published: (2025)
by: Loxley, Peter N.
Published: (2025)
Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior
by: Gros, Timo P., et al.
Published: (2025)
by: Gros, Timo P., et al.
Published: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2025)
by: Gao, Chen-Xiao, et al.
Published: (2025)
Similar Items
-
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026) -
Inferring Transition Dynamics from Value Functions
by: Adamczyk, Jacob
Published: (2025) -
Link Prediction with Untrained Message Passing Layers
by: Qarkaxhija, Lisi, et al.
Published: (2024) -
Thermodynamics of Reinforcement Learning Curricula
by: Adamczyk, Jacob, et al.
Published: (2026) -
RandomNet: Clustering Time Series Using Untrained Deep Neural Networks
by: Li, Xiaosheng, et al.
Published: (2024)