Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Xiaojun, Jiang, Xiaoguang, Li, Huiyang, Zhai, Jucai, Liu, Dengfeng, Hao, Qiaobo, Liu, Huang, Yang, Zhiguo, Xie, Ji, Gu, Ninglun, Yang, Jin, Zhang, Kailai, Bao, Yelun, Wang, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
von: Sun, Weixuan, et al.
Veröffentlicht: (2025)
von: Sun, Weixuan, et al.
Veröffentlicht: (2025)
Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
von: Liu, Dengfeng, et al.
Veröffentlicht: (2025)
von: Liu, Dengfeng, et al.
Veröffentlicht: (2025)
TN-AutoRCA: Benchmark Construction and Agentic Framework for Self-Improving Alarm-Based Root Cause Analysis in Telecommunication Networks
von: Wu, Keyu, et al.
Veröffentlicht: (2025)
von: Wu, Keyu, et al.
Veröffentlicht: (2025)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
von: Yang, Bangji, et al.
Veröffentlicht: (2026)
Fast Authenticated and Interoperable Multimedia Healthcare Data over Hybrid-Storage Blockchains
von: Yang, Jucai, et al.
Veröffentlicht: (2025)
von: Yang, Jucai, et al.
Veröffentlicht: (2025)
Theoretical Study on Structural Behaviors and Optical Properties of Silver‐Doped Silicon Anion Clusters AgSi n − ( n = 3–20)
von: Bin Liu, et al.
Veröffentlicht: (2026)
von: Bin Liu, et al.
Veröffentlicht: (2026)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
Barriers and Facilitators to the Implementation of Analgesia and Sedation Management Strategies Among ICU Healthcare Professionals: A Mixed Methods Systematic Review
von: Huiqin Shi, et al.
Veröffentlicht: (2026)
von: Huiqin Shi, et al.
Veröffentlicht: (2026)
Study on Damage Characteristics and Evolution Law of Macro‐Meso Structure of Surrounding Rock Under Mining Stress
von: Fan Lei, et al.
Veröffentlicht: (2026)
von: Fan Lei, et al.
Veröffentlicht: (2026)
The Scaling Law in Stellar Light Curves
von: Pan, Jia-Shu, et al.
Veröffentlicht: (2024)
von: Pan, Jia-Shu, et al.
Veröffentlicht: (2024)
WaveComm: Lightweight Communication for Collaborative Perception via Wavelet Feature Distillation
von: Bao, Erdemt, et al.
Veröffentlicht: (2026)
von: Bao, Erdemt, et al.
Veröffentlicht: (2026)
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving
von: Wu, Yang, et al.
Veröffentlicht: (2026)
von: Wu, Yang, et al.
Veröffentlicht: (2026)
Theoretical Study of Structures and Spectral Properties of ScSnn0/−/2− (n = 4–17) Nanoalloy Clusters
von: Shuilian Tu, et al.
Veröffentlicht: (2025)
von: Shuilian Tu, et al.
Veröffentlicht: (2025)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
von: Zheng, Bo, et al.
Veröffentlicht: (2026)
von: Zheng, Bo, et al.
Veröffentlicht: (2026)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Models
von: Xie, Guangyu, et al.
Veröffentlicht: (2025)
von: Xie, Guangyu, et al.
Veröffentlicht: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
Loss Gradient Gaussian Width based Generalization and Optimization Guarantees
von: Banerjee, Arindam, et al.
Veröffentlicht: (2024)
von: Banerjee, Arindam, et al.
Veröffentlicht: (2024)
Sketched Gaussian Mechanism for Private Federated Learning
von: Li, Qiaobo, et al.
Veröffentlicht: (2025)
von: Li, Qiaobo, et al.
Veröffentlicht: (2025)
Sketched Adaptive Federated Deep Learning: A Sharp Convergence Analysis
von: Chen, Zhijie, et al.
Veröffentlicht: (2024)
von: Chen, Zhijie, et al.
Veröffentlicht: (2024)
Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
von: Deng, Jie, et al.
Veröffentlicht: (2026)
von: Deng, Jie, et al.
Veröffentlicht: (2026)
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
Non-Hermitian higher-order topological insulators enabled by altermagnet engineering
von: Ji, Xiang, et al.
Veröffentlicht: (2025)
von: Ji, Xiang, et al.
Veröffentlicht: (2025)
In-Context Learning Distillation for Efficient Few-Shot Fine-Tuning
von: Duan, Yifei, et al.
Veröffentlicht: (2024)
von: Duan, Yifei, et al.
Veröffentlicht: (2024)
Towards Interpretable Mental Health Analysis with Large Language Models
von: Yang, Kailai, et al.
Veröffentlicht: (2023)
von: Yang, Kailai, et al.
Veröffentlicht: (2023)
Folding-Free ZNE: A Comprehensive Quantum Zero-Noise Extrapolation Approach for Mitigating Depolarizing and Decoherence Noise
von: Patil, Hrushikesh, et al.
Veröffentlicht: (2023)
von: Patil, Hrushikesh, et al.
Veröffentlicht: (2023)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning
von: Fang, Sitong, et al.
Veröffentlicht: (2025)
von: Fang, Sitong, et al.
Veröffentlicht: (2025)
Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2025)
MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
Selective Preference Optimization via Token-Level Reward Function Estimation
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
von: Yang, Kailai, et al.
Veröffentlicht: (2024)
FMDLlama: Financial Misinformation Detection based on Large Language Models
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
von: Wang, Jun, et al.
Veröffentlicht: (2025) -
CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
von: Sun, Weixuan, et al.
Veröffentlicht: (2025) -
Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
von: Liu, Dengfeng, et al.
Veröffentlicht: (2025) -
TN-AutoRCA: Benchmark Construction and Agentic Framework for Self-Improving Alarm-Based Root Cause Analysis in Telecommunication Networks
von: Wu, Keyu, et al.
Veröffentlicht: (2025) -
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)