GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Ziru, Gong, Cheng, Fu, Xinyu, Liu, Yaofang, Chen, Ran, Hu, Shoubo, Zhang, Suiyun, Liu, Rui, Zhang, Qingfu, Tu, Dandan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2025)
by: Chen, Kecheng, et al.
Published: (2025)
Efficient Reasoning via Reward Model
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
by: Chen, Jinpeng, et al.
Published: (2026)
by: Chen, Jinpeng, et al.
Published: (2026)
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2026)
by: Chen, Kecheng, et al.
Published: (2026)
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
by: Liu, Yaofang, et al.
Published: (2025)
by: Liu, Yaofang, et al.
Published: (2025)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Ultralow-dimensionality reduction for identifying critical transitions by spatial-temporal PCA
by: Chen, Pei, et al.
Published: (2025)
by: Chen, Pei, et al.
Published: (2025)
Advances in the use of organoids in endometrial diseases
by: Yaofang Liu, et al.
Published: (2024)
by: Yaofang Liu, et al.
Published: (2024)
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
by: Shan, Lianlei, et al.
Published: (2026)
by: Shan, Lianlei, et al.
Published: (2026)
LLM-Enabled Automated Algorithm Design for Multiuser Fluid Antenna Communications
by: Zheng, Gan, et al.
Published: (2026)
by: Zheng, Gan, et al.
Published: (2026)
VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
by: Chen, Rufeng, et al.
Published: (2026)
by: Chen, Rufeng, et al.
Published: (2026)
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
by: He, Jinmin, et al.
Published: (2025)
by: He, Jinmin, et al.
Published: (2025)
Adaptive Human-Computer Interaction Strategies Through Reinforcement Learning in Complex
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
LLM-Powered User Simulator for Recommender System
by: Zhang, Zijian, et al.
Published: (2024)
by: Zhang, Zijian, et al.
Published: (2024)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
Adaptive Conformal Guidance for Learning under Uncertainty
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Laser: Parameter-Efficient LLM Bi-Tuning for Sequential Recommendation with Collaborative Information
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Unify Graph Learning with Text: Unleashing LLM Potentials for Session Search
by: Wu, Songhao, et al.
Published: (2025)
by: Wu, Songhao, et al.
Published: (2025)
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models
by: Lin, Jintao, et al.
Published: (2026)
by: Lin, Jintao, et al.
Published: (2026)
LLM4AMC: Adapting Large Language Models for Adaptive Modulation and Coding
by: Pan, Xinyu, et al.
Published: (2025)
by: Pan, Xinyu, et al.
Published: (2025)
Stable Reinforcement Learning for Efficient Reasoning
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
TCR-GPT: Integrating Autoregressive Model and Reinforcement Learning for T-Cell Receptor Repertoires Generation
by: Lin, Yicheng, et al.
Published: (2024)
by: Lin, Yicheng, et al.
Published: (2024)
Scaling the Explanation of Multi-Class Bayesian Network Classifiers
by: Zhang, Yaofang, et al.
Published: (2026)
by: Zhang, Yaofang, et al.
Published: (2026)
ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving
by: Wu, Haoyuan, et al.
Published: (2025)
by: Wu, Haoyuan, et al.
Published: (2025)
Network-Adjusted Covariates for Community Detection
by: Hu, Yaofang, et al.
Published: (2023)
by: Hu, Yaofang, et al.
Published: (2023)
Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance
by: Sun, Guodong, et al.
Published: (2026)
by: Sun, Guodong, et al.
Published: (2026)
Enhancing CVRP Solver through LLM-driven Automatic Heuristic Design
by: Xie, Zhuoliang, et al.
Published: (2026)
by: Xie, Zhuoliang, et al.
Published: (2026)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
Adaptive Guidance for Local Training in Heterogeneous Federated Learning
by: Zhang, Jianqing, et al.
Published: (2024)
by: Zhang, Jianqing, et al.
Published: (2024)
Adaptive Guidance with Reinforcement Meta-Learning
by: Gaudet, Brian, et al.
Published: (2019)
by: Gaudet, Brian, et al.
Published: (2019)
Deep Generative Demand Learning for Newsvendor and Pricing
by: Gong, Shijin, et al.
Published: (2024)
by: Gong, Shijin, et al.
Published: (2024)
AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length
by: Cheng, Junhang, et al.
Published: (2025)
by: Cheng, Junhang, et al.
Published: (2025)
NavG: Risk-Aware Navigation in Crowded Environments Based on Reinforcement Learning with Guidance Points
by: Zhang, Qianyi, et al.
Published: (2025)
by: Zhang, Qianyi, et al.
Published: (2025)
Extensions of Robbins-Siegmund Theorem with Applications in Reinforcement Learning
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning
by: Liu, Weidong, et al.
Published: (2023)
by: Liu, Weidong, et al.
Published: (2023)
Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
by: Hu, Qinglong, et al.
Published: (2025)
by: Hu, Qinglong, et al.
Published: (2025)
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
by: Ming, Rui, et al.
Published: (2025)
by: Ming, Rui, et al.
Published: (2025)
Decoding Public Preferences in Volunteer Service Projects: An Empirical Study Based on Conjoint Experiment and Machine Learning Approaches
by: Rui Zhang, et al.
Published: (2025)
by: Rui Zhang, et al.
Published: (2025)
Evolve Cost-aware Acquisition Functions Using Large Language Models
by: Yao, Yiming, et al.
Published: (2024)
by: Yao, Yiming, et al.
Published: (2024)
Similar Items
-
Beyond Confidence: Adaptive and Coherent Decoding for Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2025) -
Efficient Reasoning via Reward Model
by: Wang, Yuhao, et al.
Published: (2025) -
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
by: Chen, Jinpeng, et al.
Published: (2026) -
Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models
by: Chen, Kecheng, et al.
Published: (2026) -
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
by: Liu, Yaofang, et al.
Published: (2025)