Adaptive Simulation Experiment for LLM Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Mingjie, Gao, Siyang, Hu, Jian-qiang, Zhou, Enlu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantum Grover Adaptive Search for Discrete Simulation Optimization
by: Hu, Mingjie, et al.
Published: (2026)
by: Hu, Mingjie, et al.
Published: (2026)
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
by: Wang, Xiaoshuang, et al.
Published: (2025)
by: Wang, Xiaoshuang, et al.
Published: (2025)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Online Bayesian Risk-Averse Reinforcement Learning
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Ranking and Selection with Simultaneous Input Data Collection
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
by: Li, Yingke, et al.
Published: (2026)
by: Li, Yingke, et al.
Published: (2026)
Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference
by: Li, Yingke, et al.
Published: (2026)
by: Li, Yingke, et al.
Published: (2026)
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
by: Song, Meichen, et al.
Published: (2026)
by: Song, Meichen, et al.
Published: (2026)
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning
by: Chen, Haoxuan, et al.
Published: (2026)
by: Chen, Haoxuan, et al.
Published: (2026)
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
by: Hu, Miaobo, et al.
Published: (2026)
by: Hu, Miaobo, et al.
Published: (2026)
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
by: Hu, Haichen, et al.
Published: (2026)
by: Hu, Haichen, et al.
Published: (2026)
Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning
by: Huang, Yiming, et al.
Published: (2026)
by: Huang, Yiming, et al.
Published: (2026)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Soft Adaptive Policy Optimization
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
CAWR: Corruption-Averse Advantage-Weighted Regression for Robust Policy Optimization
by: Hu, Ranting
Published: (2025)
by: Hu, Ranting
Published: (2025)
Continual Task Learning through Adaptive Policy Self-Composition
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
by: Lee, Doyeon, et al.
Published: (2026)
by: Lee, Doyeon, et al.
Published: (2026)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Optimization-Driven Adaptive Experimentation
by: Che, Ethan, et al.
Published: (2024)
by: Che, Ethan, et al.
Published: (2024)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Simulation-Based Inference for Adaptive Experiments
by: Cho, Brian M, et al.
Published: (2025)
by: Cho, Brian M, et al.
Published: (2025)
Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
by: Hu, Kai, et al.
Published: (2024)
by: Hu, Kai, et al.
Published: (2024)
On the Reliability Limits of LLM-Based Multi-Agent Planning
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
by: Pan, Chengjun, et al.
Published: (2026)
by: Pan, Chengjun, et al.
Published: (2026)
ATPO: Adaptive Tree Policy Optimization for Multi-Turn Medical Dialogue
by: Cao, Ruike, et al.
Published: (2026)
by: Cao, Ruike, et al.
Published: (2026)
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
by: Yang, Zhaohui, et al.
Published: (2025)
by: Yang, Zhaohui, et al.
Published: (2025)
Adaptive Experimental Design for Policy Learning
by: Kato, Masahiro, et al.
Published: (2024)
by: Kato, Masahiro, et al.
Published: (2024)
Robust and Efficient Zeroth-Order LLM Fine-Tuning via Adaptive Bayesian Subspace Optimizer
by: Feng, Jian, et al.
Published: (2026)
by: Feng, Jian, et al.
Published: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Exploring Grokking: Experimental and Mechanistic Investigations
by: Qiye, Hu, et al.
Published: (2024)
by: Qiye, Hu, et al.
Published: (2024)
Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting
by: Qiao, Zhongjian, et al.
Published: (2026)
by: Qiao, Zhongjian, et al.
Published: (2026)
Best Arm Identification with LLM Judges and Limited Human
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
by: Wang, Jingyao, et al.
Published: (2026)
by: Wang, Jingyao, et al.
Published: (2026)
Adaptivity and Universality: Problem-dependent Universal Regret for Online Convex Optimization
by: Zhao, Peng, et al.
Published: (2025)
by: Zhao, Peng, et al.
Published: (2025)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization
by: Lei, Kun, et al.
Published: (2023)
by: Lei, Kun, et al.
Published: (2023)
SPPCSO: Adaptive Penalized Estimation Method for High-Dimensional Correlated Data
by: Hu, Ying, et al.
Published: (2026)
by: Hu, Ying, et al.
Published: (2026)
An Adaptive Dimension Reduction Estimation Method for High-dimensional Bayesian Optimization
by: Hu, Shouri, et al.
Published: (2024)
by: Hu, Shouri, et al.
Published: (2024)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
by: Hu, Xixi, et al.
Published: (2024)
by: Hu, Xixi, et al.
Published: (2024)
Similar Items
-
Quantum Grover Adaptive Search for Discrete Simulation Optimization
by: Hu, Mingjie, et al.
Published: (2026) -
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
by: Wang, Xiaoshuang, et al.
Published: (2025) -
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024) -
Online Bayesian Risk-Averse Reinforcement Learning
by: Wang, Yuhao, et al.
Published: (2025) -
Ranking and Selection with Simultaneous Input Data Collection
by: Wang, Yuhao, et al.
Published: (2025)