Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhaochun, Wang, Chen, Bai, Jionghao, Cui, Shisheng, Lan, Ge, Zhao, Zhou, Wang, Yue |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Towards a Theoretical Understanding to the Generalization of RLHF
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
di: Li, Zhaochun, et al.
Pubblicazione: (2026)
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
di: Li, Xiaofan, et al.
Pubblicazione: (2026)
di: Li, Xiaofan, et al.
Pubblicazione: (2026)
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
di: Wei, Wang, et al.
Pubblicazione: (2025)
di: Wei, Wang, et al.
Pubblicazione: (2025)
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
di: Song, Meichen, et al.
Pubblicazione: (2026)
di: Song, Meichen, et al.
Pubblicazione: (2026)
On the Trade-off between Flatness and Optimization in Distributed Learning
di: Cao, Ying, et al.
Pubblicazione: (2024)
di: Cao, Ying, et al.
Pubblicazione: (2024)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
di: Yao, Xincheng, et al.
Pubblicazione: (2026)
di: Yao, Xincheng, et al.
Pubblicazione: (2026)
ImprovDML: Improved Trade-off in Private Byzantine-Resilient Distributed Machine Learning
di: Liu, Bing, et al.
Pubblicazione: (2025)
di: Liu, Bing, et al.
Pubblicazione: (2025)
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs
di: Norman, Ben, et al.
Pubblicazione: (2023)
di: Norman, Ben, et al.
Pubblicazione: (2023)
Characterizing the Accuracy-Communication-Privacy Trade-off in Distributed Stochastic Convex Optimization
di: Salgia, Sudeep, et al.
Pubblicazione: (2025)
di: Salgia, Sudeep, et al.
Pubblicazione: (2025)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
di: Su, Xuerui, et al.
Pubblicazione: (2025)
di: Su, Xuerui, et al.
Pubblicazione: (2025)
Neural Exploitation and Exploration of Contextual Bandits
di: Ban, Yikun, et al.
Pubblicazione: (2023)
di: Ban, Yikun, et al.
Pubblicazione: (2023)
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training
di: Wang, Chen, et al.
Pubblicazione: (2026)
di: Wang, Chen, et al.
Pubblicazione: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
di: Alahmadi, Muhammad J., et al.
Pubblicazione: (2026)
di: Alahmadi, Muhammad J., et al.
Pubblicazione: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
di: Liang, Kun, et al.
Pubblicazione: (2026)
di: Liang, Kun, et al.
Pubblicazione: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
di: Chen, Zhipeng, et al.
Pubblicazione: (2025)
di: Chen, Zhipeng, et al.
Pubblicazione: (2025)
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
di: Yan, Renye, et al.
Pubblicazione: (2024)
di: Yan, Renye, et al.
Pubblicazione: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
di: Dai, Zhenwen, et al.
Pubblicazione: (2024)
di: Dai, Zhenwen, et al.
Pubblicazione: (2024)
Exploitation Is All You Need... for Exploration
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
di: Rentschler, Micah, et al.
Pubblicazione: (2025)
Provably Efficient Exploration in Policy Optimization
di: Cai, Qi, et al.
Pubblicazione: (2019)
di: Cai, Qi, et al.
Pubblicazione: (2019)
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
di: Bamba, Udbhav, et al.
Pubblicazione: (2025)
di: Bamba, Udbhav, et al.
Pubblicazione: (2025)
Controllable Pareto Trade-off between Fairness and Accuracy
di: Du, Yongkang, et al.
Pubblicazione: (2025)
di: Du, Yongkang, et al.
Pubblicazione: (2025)
Architecture Selection via the Trade-off Between Accuracy and Robustness
di: Deng, Zhun, et al.
Pubblicazione: (2019)
di: Deng, Zhun, et al.
Pubblicazione: (2019)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
di: Zeng, Weihao, et al.
Pubblicazione: (2024)
di: Zeng, Weihao, et al.
Pubblicazione: (2024)
Exploration-Exploitation Tradeoff in Universal Lossy Compression
di: Weinberger, Nir, et al.
Pubblicazione: (2025)
di: Weinberger, Nir, et al.
Pubblicazione: (2025)
Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action
di: Gao, Gong, et al.
Pubblicazione: (2026)
di: Gao, Gong, et al.
Pubblicazione: (2026)
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
di: Zhao, Yang, et al.
Pubblicazione: (2026)
di: Zhao, Yang, et al.
Pubblicazione: (2026)
Private Optimal Inventory Policy Learning for Feature-based Newsvendor with Unknown Demand
di: Zhao, Tuoyi, et al.
Pubblicazione: (2024)
di: Zhao, Tuoyi, et al.
Pubblicazione: (2024)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
di: K, Swaminathan S, et al.
Pubblicazione: (2026)
di: K, Swaminathan S, et al.
Pubblicazione: (2026)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
di: Wang, Yue, et al.
Pubblicazione: (2025)
di: Wang, Yue, et al.
Pubblicazione: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
Learning What Matters: Adaptive Information-Theoretic Objectives for Robot Exploration
di: Yu, Youwei, et al.
Pubblicazione: (2026)
di: Yu, Youwei, et al.
Pubblicazione: (2026)
Learning Against Distributional Uncertainty: On the Trade-off Between Robustness and Specificity
di: Wang, Shixiong, et al.
Pubblicazione: (2023)
di: Wang, Shixiong, et al.
Pubblicazione: (2023)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
di: Chen, Lin, et al.
Pubblicazione: (2026)
di: Chen, Lin, et al.
Pubblicazione: (2026)
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
di: Han, Xiaotian, et al.
Pubblicazione: (2025)
di: Han, Xiaotian, et al.
Pubblicazione: (2025)
Enhancing Trade-offs in Privacy, Utility, and Computational Efficiency through MUltistage Sampling Technique (MUST)
di: Zhao, Xingyuan, et al.
Pubblicazione: (2023)
di: Zhao, Xingyuan, et al.
Pubblicazione: (2023)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
di: Lam, Phong, et al.
Pubblicazione: (2026)
di: Lam, Phong, et al.
Pubblicazione: (2026)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
di: Kim, JaeYoon, et al.
Pubblicazione: (2024)
di: Kim, JaeYoon, et al.
Pubblicazione: (2024)
Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2026)
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
di: Wang, Chen, et al.
Pubblicazione: (2025) -
Towards a Theoretical Understanding to the Generalization of RLHF
di: Li, Zhaochun, et al.
Pubblicazione: (2026) -
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
di: Li, Xiaofan, et al.
Pubblicazione: (2026) -
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
di: Wei, Wang, et al.
Pubblicazione: (2025) -
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
di: Song, Meichen, et al.
Pubblicazione: (2026)