In-context Exploration-Exploitation for Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Zhenwen, Tomasi, Federico, Ghiassian, Sina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning in complex action spaces without policy gradients
by: Tavakoli, Arash, et al.
Published: (2024)
by: Tavakoli, Arash, et al.
Published: (2024)
Soft Preference Optimization: Aligning Language Models to Expert Distributions
by: Sharifnassab, Arsalan, et al.
Published: (2024)
by: Sharifnassab, Arsalan, et al.
Published: (2024)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
by: Liu, Jia, et al.
Published: (2025)
by: Liu, Jia, et al.
Published: (2025)
Auxiliary task discovery through generate-and-test
by: Rafiee, Banafsheh, et al.
Published: (2022)
by: Rafiee, Banafsheh, et al.
Published: (2022)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
by: Dai, Runpeng, et al.
Published: (2025)
by: Dai, Runpeng, et al.
Published: (2025)
Adaptive Data Exploitation in Deep Reinforcement Learning
by: Yuan, Mingqi, et al.
Published: (2025)
by: Yuan, Mingqi, et al.
Published: (2025)
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs
by: Norman, Ben, et al.
Published: (2023)
by: Norman, Ben, et al.
Published: (2023)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
by: Lam, Phong, et al.
Published: (2026)
by: Lam, Phong, et al.
Published: (2026)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
by: Mukherjee, Sumantrak, et al.
Published: (2025)
by: Mukherjee, Sumantrak, et al.
Published: (2025)
Disentangling Exploration of Large Language Models by Optimal Exploitation
by: Grams, Tim, et al.
Published: (2025)
by: Grams, Tim, et al.
Published: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
by: Panaganti, Kishan, et al.
Published: (2026)
by: Panaganti, Kishan, et al.
Published: (2026)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Satisficing Exploration for Deep Reinforcement Learning
by: Arumugam, Dilip, et al.
Published: (2024)
by: Arumugam, Dilip, et al.
Published: (2024)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Finding Optimal Trading History in Reinforcement Learning for Stock Market Trading
by: Montazeri, Sina, et al.
Published: (2025)
by: Montazeri, Sina, et al.
Published: (2025)
REFN: A Reinforcement-Learning-From-Network Framework against 1-day/n-day Exploitations
by: Yu, Tianlong, et al.
Published: (2025)
by: Yu, Tianlong, et al.
Published: (2025)
Variable-Agnostic Causal Exploration for Reinforcement Learning
by: Nguyen, Minh Hoang, et al.
Published: (2024)
by: Nguyen, Minh Hoang, et al.
Published: (2024)
Exploration in Knowledge Transfer Utilizing Reinforcement Learning
by: Jedlička, Adam, et al.
Published: (2024)
by: Jedlička, Adam, et al.
Published: (2024)
Neighboring State-based Exploration for Reinforcement Learning
by: Li, Yu-Teng, et al.
Published: (2022)
by: Li, Yu-Teng, et al.
Published: (2022)
Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
by: Berseth, Glen
Published: (2025)
by: Berseth, Glen
Published: (2025)
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
by: Balloch, Jonathan C., et al.
Published: (2024)
by: Balloch, Jonathan C., et al.
Published: (2024)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
Enhance Exploration in Safe Reinforcement Learning with Contrastive Representation Learning
by: Doan, Duc Kien, et al.
Published: (2025)
by: Doan, Duc Kien, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
Offline Model-Based Reinforcement Learning with Anti-Exploration
by: Srinivasan, Padmanaba, et al.
Published: (2024)
by: Srinivasan, Padmanaba, et al.
Published: (2024)
A Temporally Correlated Latent Exploration for Reinforcement Learning
by: Oh, SuMin, et al.
Published: (2024)
by: Oh, SuMin, et al.
Published: (2024)
Adventurer: Exploration with BiGAN for Deep Reinforcement Learning
by: Liu, Yongshuai, et al.
Published: (2025)
by: Liu, Yongshuai, et al.
Published: (2025)
Guardian: Decoupling Exploration from Safety in Reinforcement Learning
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
by: Xu, Fangzhi, et al.
Published: (2025)
by: Xu, Fangzhi, et al.
Published: (2025)
Similar Items
-
Learning in complex action spaces without policy gradients
by: Tavakoli, Arash, et al.
Published: (2024) -
Soft Preference Optimization: Aligning Language Models to Expert Distributions
by: Sharifnassab, Arsalan, et al.
Published: (2024) -
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025) -
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
by: Liu, Jia, et al.
Published: (2025) -
Auxiliary task discovery through generate-and-test
by: Rafiee, Banafsheh, et al.
Published: (2022)