EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Siyao, Ma, Cong, Cheng, Zhihao, Lei, Shiye, Li, Minghao, Zeng, Ying, Tou, Huaixiao, Jia, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
by: Lei, Shiye, et al.
Published: (2026)
by: Lei, Shiye, et al.
Published: (2026)
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
by: Li, Minghao, et al.
Published: (2025)
by: Li, Minghao, et al.
Published: (2025)
Revisiting LLM Reasoning via Information Bottleneck
by: Lei, Shiye, et al.
Published: (2025)
by: Lei, Shiye, et al.
Published: (2025)
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA
by: Zeng, Yunsheng, et al.
Published: (2026)
by: Zeng, Yunsheng, et al.
Published: (2026)
Offline Behavioral Data Selection
by: Lei, Shiye, et al.
Published: (2025)
by: Lei, Shiye, et al.
Published: (2025)
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
by: Li, Zheng, et al.
Published: (2026)
by: Li, Zheng, et al.
Published: (2026)
MemRerank: Preference Memory for Personalized Product Reranking
by: Peng, Zhiyuan, et al.
Published: (2026)
by: Peng, Zhiyuan, et al.
Published: (2026)
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
by: Tou, Huaixiao, et al.
Published: (2025)
by: Tou, Huaixiao, et al.
Published: (2025)
Offline Behavior Distillation
by: Lei, Shiye, et al.
Published: (2024)
by: Lei, Shiye, et al.
Published: (2024)
TRAWL: External Knowledge-Enhanced Recommendation with LLM Assistance
by: Luo, Weiqing, et al.
Published: (2024)
by: Luo, Weiqing, et al.
Published: (2024)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation
by: Wang, Zezhou, et al.
Published: (2026)
by: Wang, Zezhou, et al.
Published: (2026)
Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
by: Ma, Hao, et al.
Published: (2025)
by: Ma, Hao, et al.
Published: (2025)
Relative Policy-Transition Optimization for Fast Policy Transfer
by: Xu, Jiawei, et al.
Published: (2022)
by: Xu, Jiawei, et al.
Published: (2022)
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
Drift-aware Collaborative Assistance Mixture of Experts for Heterogeneous Multistream Learning
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning
by: Yan, Zhengqing, et al.
Published: (2026)
by: Yan, Zhengqing, et al.
Published: (2026)
Agents on a Tree: Pathwise Coordination for Multi-Objective Molecular Optimization
by: Zhang, Jia, et al.
Published: (2026)
by: Zhang, Jia, et al.
Published: (2026)
Adaptive Masking Enhances Visual Grounding
by: Jia, Sen, et al.
Published: (2024)
by: Jia, Sen, et al.
Published: (2024)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
by: Gao, Lei, et al.
Published: (2026)
by: Gao, Lei, et al.
Published: (2026)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
Mixture-of-Schedulers: An Adaptive Scheduling Agent as a Learned Router for Expert Policies
by: Wang, Xinbo, et al.
Published: (2025)
by: Wang, Xinbo, et al.
Published: (2025)
Phenotypic Profile-Informed Generation of Drug-Like Molecules via Dual-Channel Variational Autoencoders
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Optimus-3: Dual-Router Aligned Mixture-of-Experts Agent with Dual-Granularity Reasoning-Aware Policy Optimization
by: Li, Zaijing, et al.
Published: (2025)
by: Li, Zaijing, et al.
Published: (2025)
Difficulty-Estimated Policy Optimization
by: Zhao, Yu, et al.
Published: (2026)
by: Zhao, Yu, et al.
Published: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
by: Wang, Jingyi, et al.
Published: (2026)
by: Wang, Jingyi, et al.
Published: (2026)
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
by: He, Shuo, et al.
Published: (2026)
by: He, Shuo, et al.
Published: (2026)
Truncated Proximal Policy Optimization
by: Fan, Tiantian, et al.
Published: (2025)
by: Fan, Tiantian, et al.
Published: (2025)
TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
by: Ma, Shichao, et al.
Published: (2026)
by: Ma, Shichao, et al.
Published: (2026)
L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts
by: Yang, Minghao, et al.
Published: (2026)
by: Yang, Minghao, et al.
Published: (2026)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
tcrLM: a lightweight protein language model for predicting T cell receptor and epitope binding specificity
by: Fang, Xing, et al.
Published: (2024)
by: Fang, Xing, et al.
Published: (2024)
Graph-Enhanced Policy Optimization in LLM Agent Training
by: Yuan, Jiazhen, et al.
Published: (2025)
by: Yuan, Jiazhen, et al.
Published: (2025)
Large Language Models Need Consultants for Reasoning: Becoming an Expert in a Complex Human System Through Behavior Simulation
by: Wang, Chuwen, et al.
Published: (2024)
by: Wang, Chuwen, et al.
Published: (2024)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
by: Li, Dengchun, et al.
Published: (2024)
by: Li, Dengchun, et al.
Published: (2024)
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
by: He, Shuo, et al.
Published: (2026)
by: He, Shuo, et al.
Published: (2026)
Similar Items
-
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
by: Lei, Shiye, et al.
Published: (2026) -
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
by: Li, Minghao, et al.
Published: (2025) -
Revisiting LLM Reasoning via Information Bottleneck
by: Lei, Shiye, et al.
Published: (2025) -
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA
by: Zeng, Yunsheng, et al.
Published: (2026) -
Offline Behavioral Data Selection
by: Lei, Shiye, et al.
Published: (2025)