DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Xiaoyu, Zhao, Sitong, Wang, Haotian, Chen, Shuaiting, Peng, Yiping, Ji, Yunjie, Zhao, Han, Li, Xiangang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
by: Ji, Yunjie, et al.
Published: (2025)
by: Ji, Yunjie, et al.
Published: (2025)
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
by: Wang, Haotian, et al.
Published: (2025)
by: Wang, Haotian, et al.
Published: (2025)
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
by: Ji, Yunjie, et al.
Published: (2025)
by: Ji, Yunjie, et al.
Published: (2025)
Not All Correct Answers Are Equal: Why Your Distillation Source Matters
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
by: Xu, Ruoxi, et al.
Published: (2025)
by: Xu, Ruoxi, et al.
Published: (2025)
Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training
by: Cao, Chuxue, et al.
Published: (2026)
by: Cao, Chuxue, et al.
Published: (2026)
Unsupervised Deep Equilibrium Model Learning for Large-Scale Channel Estimation with Performance Guarantees
by: Tian, Haotian, et al.
Published: (2025)
by: Tian, Haotian, et al.
Published: (2025)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
On the Difficulty of Learning a Meta-network for Training Data Selection
by: Du, Zilin, et al.
Published: (2026)
by: Du, Zilin, et al.
Published: (2026)
Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models
by: Lu, Li, et al.
Published: (2026)
by: Lu, Li, et al.
Published: (2026)
The "Graded Difficulty" Library
by: Geeslin, Robert H.
Published: (1971)
by: Geeslin, Robert H.
Published: (1971)
Distilling LLM Reasoning into Graph of Concept Predictors
by: Yu, Ziyang, et al.
Published: (2026)
by: Yu, Ziyang, et al.
Published: (2026)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Knowledge from Large-Scale Protein Contact Prediction Models Can Be Transferred to the Data-Scarce RNA Contact Prediction Task
by: Jian, Yiren, et al.
Published: (2023)
by: Jian, Yiren, et al.
Published: (2023)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025)
by: Kim, Minwu, et al.
Published: (2025)
ZPD Detector: Data Selection via Capability-Difficulty Alignment for Large Language Models
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
by: Su, DiJia, et al.
Published: (2025)
by: Su, DiJia, et al.
Published: (2025)
LLM-based Privacy Data Augmentation Guided by Knowledge Distillation with a Distribution Tutor for Medical Text Classification
by: Song, Yiping, et al.
Published: (2024)
by: Song, Yiping, et al.
Published: (2024)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading
by: Yu, Qianjin, et al.
Published: (2025)
by: Yu, Qianjin, et al.
Published: (2025)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
Enhancing Multi-Hop Knowledge Graph Reasoning through Reward Shaping Techniques
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
by: Zhao, Zhengyang, et al.
Published: (2026)
by: Zhao, Zhengyang, et al.
Published: (2026)
Fully Distributed State Estimation for Multi-agent Systems and its Application in Cooperative Localization
by: Huang, Shuaiting, et al.
Published: (2025)
by: Huang, Shuaiting, et al.
Published: (2025)
Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation
by: He, Xiao, et al.
Published: (2026)
by: He, Xiao, et al.
Published: (2026)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
by: Wang, Junda, et al.
Published: (2024)
by: Wang, Junda, et al.
Published: (2024)
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
by: Li, Tenghui, et al.
Published: (2024)
by: Li, Tenghui, et al.
Published: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
On Data Engineering for Scaling LLM Terminal Capabilities
by: Pi, Renjie, et al.
Published: (2026)
by: Pi, Renjie, et al.
Published: (2026)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models
by: Xiaoyu, Wang, et al.
Published: (2026)
by: Xiaoyu, Wang, et al.
Published: (2026)
Similar Items
-
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
by: Ji, Yunjie, et al.
Published: (2025) -
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
by: Wang, Haotian, et al.
Published: (2025) -
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
by: Zhao, Han, et al.
Published: (2025) -
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
by: Tian, Xiaoyu, et al.
Published: (2025) -
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
by: Ji, Yunjie, et al.
Published: (2025)