Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Chen, Ma, Yiyuan, Yang, Yuan, Liu, Deyi, Liu, Jing, Song, Zuquan, Song, Yuxin, Ren, Cheng, Zhu, Hang, Liu, Xin, Qiao, Siyuan, Zhou, Xun, Xiang, Liang, Wu, Yonghui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025)
by: Zheng, Chen, et al.
Published: (2025)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
by: Lv, Ang, et al.
Published: (2025)
by: Lv, Ang, et al.
Published: (2025)
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
by: Yu, Qifan, et al.
Published: (2026)
by: Yu, Qifan, et al.
Published: (2026)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in Large Language Models
by: Li, Yiyuan, et al.
Published: (2024)
by: Li, Yiyuan, et al.
Published: (2024)
Regulating Cocatalyst Spin‐Electronic Structures for Achieving Solar‐To‐H 2 of 7.32% in Photothermal‐Catalytic H 2 O Overall Splitting
by: Yiyuan Liu, et al.
Published: (2025)
by: Yiyuan Liu, et al.
Published: (2025)
Reinforcement Learning Method for Zero-Sum Linear-Quadratic Stochastic Differential Games in Infinite Horizons
by: Wang, Yiyuan
Published: (2026)
by: Wang, Yiyuan
Published: (2026)
A New Algorithm for Computing the Stabilizing Solution of General Periodic Time-Varying Stochastic Game-Theoretic Riccati Differential Equations
by: Wang, Yiyuan
Published: (2025)
by: Wang, Yiyuan
Published: (2025)
A Unified Computational Approach for Zero-Sum Linear-Quadratic Stochastic Differential Games in Infinite Horizons
by: Wang, Yiyuan
Published: (2025)
by: Wang, Yiyuan
Published: (2025)
A Convergent Algorithm Based on Deterministic Approximation for a Large Class of Regime-Switching Generalized Stochastic Game-Theoretic Riccati Differential Equations
by: Wang, Yiyuan
Published: (2025)
by: Wang, Yiyuan
Published: (2025)
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs
by: Zheng, Chen, et al.
Published: (2024)
by: Zheng, Chen, et al.
Published: (2024)
Model Merging in Pre-training of Large Language Models
by: Li, Yunshui, et al.
Published: (2025)
by: Li, Yunshui, et al.
Published: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF
by: Liu, Jing
Published: (2025)
by: Liu, Jing
Published: (2025)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
by: Zheng, Chen, et al.
Published: (2024)
by: Zheng, Chen, et al.
Published: (2024)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Subgroup learning in functional regression models under the RKHS framework
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
Change-plane analysis in functional response quantile regression
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration
by: Pan, Yiyuan, et al.
Published: (2025)
by: Pan, Yiyuan, et al.
Published: (2025)
Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
by: Xu, Yunzhe, et al.
Published: (2025)
by: Xu, Yunzhe, et al.
Published: (2025)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
A Novel Framework for Modeling Quarantinable Disease Transmission
by: Liu, Wenchen, et al.
Published: (2025)
by: Liu, Wenchen, et al.
Published: (2025)
Segment-Level Attribution for Selective Learning of Long Reasoning Traces
by: Wang, Siyuan, et al.
Published: (2026)
by: Wang, Siyuan, et al.
Published: (2026)
Mechanisms promoting biodiversity in ecosystems
by: Kang, Ju, et al.
Published: (2024)
by: Kang, Ju, et al.
Published: (2024)
BEMEval-Doc2Schema: Benchmarking Large Language Models for Structured Data Extraction in Building Energy Modeling
by: Jia, Yiyuan, et al.
Published: (2026)
by: Jia, Yiyuan, et al.
Published: (2026)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
by: Li, Handong, et al.
Published: (2025)
by: Li, Handong, et al.
Published: (2025)
Co-Scheduling of Energy and Production in Discrete Manufacturing Considering Decision-Dependent Uncertainties
by: Pan, Yiyuan, et al.
Published: (2024)
by: Pan, Yiyuan, et al.
Published: (2024)
Improving log-based anomaly detection through learned adaptive filter
by: Xiong, Yiyuan, et al.
Published: (2025)
by: Xiong, Yiyuan, et al.
Published: (2025)
Deep Reinforcement Learning for Artificial Upwelling Energy Management
by: Zhang, Yiyuan, et al.
Published: (2023)
by: Zhang, Yiyuan, et al.
Published: (2023)
MambaUIE&SR: Unraveling the Ocean's Secrets with Only 2.8 GFLOPs
by: Chen, Zhihao, et al.
Published: (2024)
by: Chen, Zhihao, et al.
Published: (2024)
Part-Attention Based Model Make Occluded Person Re-Identification Stronger
by: Chen, Zhihao, et al.
Published: (2024)
by: Chen, Zhihao, et al.
Published: (2024)
Mind the Jumps: A Scalable Robust Local Gaussian Process for Multidimensional Response Surfaces with Discontinuities
by: Adjetey, Isaac, et al.
Published: (2025)
by: Adjetey, Isaac, et al.
Published: (2025)
Trajectories of temperamental shyness and attention shifting in Chinese children and their relations to psychosocial adjustment at entry to elementary school
by: Yiyuan Xu, et al.
Published: (2024)
by: Yiyuan Xu, et al.
Published: (2024)
Efficient Federated RLHF via Zeroth-Order Policy Optimization
by: Wang, Deyi, et al.
Published: (2026)
by: Wang, Deyi, et al.
Published: (2026)
FLAME: Learning to Navigate with Multimodal LLM in Urban Environments
by: Xu, Yunzhe, et al.
Published: (2024)
by: Xu, Yunzhe, et al.
Published: (2024)
Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation
by: Pan, Yiyuan, et al.
Published: (2024)
by: Pan, Yiyuan, et al.
Published: (2024)
Diversified Scaling Inference in Time Series Foundation Models
by: Hua, Ruijin, et al.
Published: (2026)
by: Hua, Ruijin, et al.
Published: (2026)
Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation
by: Pan, Yiyuan, et al.
Published: (2025)
by: Pan, Yiyuan, et al.
Published: (2025)
Time-RA: Towards Time Series Reasoning for Anomaly Diagnosis with LLM Feedback
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
by: Liang, Jiawei, et al.
Published: (2022)
by: Liang, Jiawei, et al.
Published: (2022)
Similar Items
-
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025) -
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
by: Lv, Ang, et al.
Published: (2025) -
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
by: Yu, Qifan, et al.
Published: (2026) -
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026) -
FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in Large Language Models
by: Li, Yiyuan, et al.
Published: (2024)