Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Yuelin, Cheng, Zhengxue, Liu, Wei, Song, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
von: Wen, Muning, et al.
Veröffentlicht: (2024)
von: Wen, Muning, et al.
Veröffentlicht: (2024)
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
von: Choo, Jinho, et al.
Veröffentlicht: (2026)
von: Choo, Jinho, et al.
Veröffentlicht: (2026)
Training Large Language Models to Reason via EM Policy Gradient
von: Xu, Tianbing
Veröffentlicht: (2025)
von: Xu, Tianbing
Veröffentlicht: (2025)
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
MeetBench-XL: Calibrated Multi-Dimensional Evaluation and Learned Dual-Policy Agents for Real-Time Meetings
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
Scalable Token-Level Hallucination Detection in Large Language Models
von: Min, Rui, et al.
Veröffentlicht: (2026)
von: Min, Rui, et al.
Veröffentlicht: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
ProToken: Token-Level Attribution for Federated Large Language Models
von: Gill, Waris, et al.
Veröffentlicht: (2026)
von: Gill, Waris, et al.
Veröffentlicht: (2026)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
von: Hu, Haoyu, et al.
Veröffentlicht: (2026)
von: Hu, Haoyu, et al.
Veröffentlicht: (2026)
GIO: Gradient Information Optimization for Training Dataset Selection
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Backward-Friendly Optimization: Training Large Language Models with Approximate Gradients under Memory Constraints
von: Yang, Jing, et al.
Veröffentlicht: (2025)
von: Yang, Jing, et al.
Veröffentlicht: (2025)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025)
von: Lan, Disen, et al.
Veröffentlicht: (2025)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
von: Lee, Deokjae, et al.
Veröffentlicht: (2024)
von: Lee, Deokjae, et al.
Veröffentlicht: (2024)
Policy-Gradient Training of Language Models for Ranking
von: Gao, Ge, et al.
Veröffentlicht: (2023)
von: Gao, Ge, et al.
Veröffentlicht: (2023)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
von: Guo, Yiran, et al.
Veröffentlicht: (2025)
von: Guo, Yiran, et al.
Veröffentlicht: (2025)
Selective Preference Optimization via Token-Level Reward Function Estimation
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
von: Li, Chengao, et al.
Veröffentlicht: (2025)
von: Li, Chengao, et al.
Veröffentlicht: (2025)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
von: Fang, Yangyi, et al.
Veröffentlicht: (2026)
von: Fang, Yangyi, et al.
Veröffentlicht: (2026)
Cross-Entropy Optimization for Hyperparameter Optimization in Stochastic Gradient-based Approaches to Train Deep Neural Networks
von: Li, Kevin, et al.
Veröffentlicht: (2024)
von: Li, Kevin, et al.
Veröffentlicht: (2024)
Large Language Models Struggle in Token-Level Clinical Named Entity Recognition
von: Lu, Qiuhao, et al.
Veröffentlicht: (2024)
von: Lu, Qiuhao, et al.
Veröffentlicht: (2024)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiyang, et al.
Veröffentlicht: (2026)
ESPO: Entropy Importance Sampling Policy Optimization
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Promising Tokens for Large Language Models
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2026)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2026)
Training-Trajectory-Aware Token Selection
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
von: Kesgin, H. Toprak, et al.
Veröffentlicht: (2024)
von: Kesgin, H. Toprak, et al.
Veröffentlicht: (2024)
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2025)
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
Provable Training Data Identification for Large Language Models
von: Liu, Zhenlong, et al.
Veröffentlicht: (2025)
von: Liu, Zhenlong, et al.
Veröffentlicht: (2025)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
Autoregressive Policy Optimization for Constrained Allocation Tasks
von: Winkel, David, et al.
Veröffentlicht: (2024)
von: Winkel, David, et al.
Veröffentlicht: (2024)
Policy Gradient with Adaptive Entropy Annealing for Continual Fine-Tuning
von: Zhang, Yaqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yaqian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026) -
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026) -
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026) -
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
von: Wen, Muning, et al.
Veröffentlicht: (2024) -
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
von: Choo, Jinho, et al.
Veröffentlicht: (2026)