SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Binbin, Ma, Xing, Liang, Yiheng, Ruan, Jingqing, Fu, Xiaoliang, Lin, Kepeng, Zhu, Benchang, Zeng, Ke, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
by: Yang, Haojin, et al.
Published: (2026)
by: Yang, Haojin, et al.
Published: (2026)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
by: Jian, Ai, et al.
Published: (2026)
by: Jian, Ai, et al.
Published: (2026)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
by: Jia, Nan, et al.
Published: (2026)
by: Jia, Nan, et al.
Published: (2026)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
Explainable Reinforcement Learning via a Causal World Model
by: Yu, Zhongwei, et al.
Published: (2023)
by: Yu, Zhongwei, et al.
Published: (2023)
Learning Causal Dynamics Models in Object-Oriented Environments
by: Yu, Zhongwei, et al.
Published: (2024)
by: Yu, Zhongwei, et al.
Published: (2024)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
by: Xie, Xinyu, et al.
Published: (2025)
by: Xie, Xinyu, et al.
Published: (2025)
Is Market a Hero or Villain? A Narrative Policy Framework Analysis of China's Patient‐Centred Healthcare Policy
by: Jingqing Yang
Published: (2026)
by: Jingqing Yang
Published: (2026)
Multitask Vehicle Signal Recognition With Dual‐Speed Adaptive Weighting
by: Dianjing Cheng, et al.
Published: (2025)
by: Dianjing Cheng, et al.
Published: (2025)
Decision-Path Patterns as Tree Reliability Signals: Path-based Adaptive Weighting for Random Forest Classification
by: Park, Youngjoon
Published: (2026)
by: Park, Youngjoon
Published: (2026)
Learning Top-k Subtask Planning Tree based on Discriminative Representation Pre-training for Decision Making
by: Ruan, Jingqing, et al.
Published: (2023)
by: Ruan, Jingqing, et al.
Published: (2023)
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
SCOPE in Cataloguing.
by: Tom, Ellen, et al.
Published: (1970)
by: Tom, Ellen, et al.
Published: (1970)
Stability Evaluation of the Goaf Based on Combination Weighting and Cloud Model
by: Linning Guo, et al.
Published: (2024)
by: Linning Guo, et al.
Published: (2024)
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
by: Cheng, Jiaming, et al.
Published: (2025)
by: Cheng, Jiaming, et al.
Published: (2025)
Knowledge Distillation for Variational Quantum Convolutional Neural Networks on Heterogeneous Data
by: Yu, Kai, et al.
Published: (2025)
by: Yu, Kai, et al.
Published: (2025)
Data-Efficient On-Policy Distillation for Automatic Speech Recognition
by: Lin, Yu, et al.
Published: (2026)
by: Lin, Yu, et al.
Published: (2026)
Generative AI Driven Task-Oriented Adaptive Semantic Communications
by: Fu, Yuzhou, et al.
Published: (2024)
by: Fu, Yuzhou, et al.
Published: (2024)
Dual-Path Enhancements in Event-Based Eye Tracking: Augmented Robustness and Adaptive Temporal Modeling
by: Truong, Hoang M., et al.
Published: (2025)
by: Truong, Hoang M., et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
A Center‐Border Dual‐Branch Network With Dynamic Weighted Fusion for Breast Cancer Histopathology Image Classification
by: Liang Zeng, et al.
Published: (2026)
by: Liang Zeng, et al.
Published: (2026)
CBNN: 3-Party Secure Framework for Customized Binary Neural Networks Inference
by: Dong, Benchang, et al.
Published: (2024)
by: Dong, Benchang, et al.
Published: (2024)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
by: Kwan, Wai-Chung, et al.
Published: (2026)
by: Kwan, Wai-Chung, et al.
Published: (2026)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
by: Yuan, Xihao, et al.
Published: (2025)
by: Yuan, Xihao, et al.
Published: (2025)
CoSLight: Co-optimizing Collaborator Selection and Decision-making to Enhance Traffic Signal Control
by: Ruan, Jingqing, et al.
Published: (2024)
by: Ruan, Jingqing, et al.
Published: (2024)
Distilling Reasoning Ability from Large Language Models with Adaptive Thinking
by: Chen, Xiaoshu, et al.
Published: (2024)
by: Chen, Xiaoshu, et al.
Published: (2024)
Two-layer consensus based on master-slave consortium chain data sharing for Internet of Vehicles
by: Zhao, Feng, et al.
Published: (2024)
by: Zhao, Feng, et al.
Published: (2024)
SCOPE-DTI: Semi-Inductive Dataset Construction and Framework Optimization for Practical Usability Enhancement in Deep Learning-Based Drug Target Interaction Prediction
by: Chen, Yigang, et al.
Published: (2025)
by: Chen, Yigang, et al.
Published: (2025)
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
by: Ren, Tianfei, et al.
Published: (2026)
by: Ren, Tianfei, et al.
Published: (2026)
HackaLOD: HAICu - SCOPE
by: Romein, C. A. (Annemieke)
Published: (2025)
by: Romein, C. A. (Annemieke)
Published: (2025)
SCOPE for Hexapod Gait Generation
by: O'Connor, Jim, et al.
Published: (2025)
by: O'Connor, Jim, et al.
Published: (2025)
Similar Items
-
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
by: Jian, Ai, et al.
Published: (2025) -
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
by: Zhang, Xiaoyun, et al.
Published: (2025) -
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
by: Yang, Haojin, et al.
Published: (2026) -
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
by: Zhang, Xiaoyun, et al.
Published: (2025) -
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026)