DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chengyu, Yan, Junbing, Yue, Yuanhao, Huang, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
by: Lyu, Yuanjie, et al.
Published: (2026)
by: Lyu, Yuanjie, et al.
Published: (2026)
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)
by: Yue, Yuanhao, et al.
Published: (2026)
Do Large Language Models Understand Logic or Just Mimick Context?
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
From Correction to Mastery: Reinforced Distillation of Large Language Model Agents
by: Lyu, Yuanjie, et al.
Published: (2025)
by: Lyu, Yuanjie, et al.
Published: (2025)
Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
Qwen2.5 Technical Report
by: Qwen, et al.
Published: (2024)
by: Qwen, et al.
Published: (2024)
TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language Models
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Qwen2.5-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025)
by: Bai, Shuai, et al.
Published: (2025)
Qwen2.5-Coder Technical Report
by: Hui, Binyuan, et al.
Published: (2024)
by: Hui, Binyuan, et al.
Published: (2024)
Qwen2.5-1M Technical Report
by: Yang, An, et al.
Published: (2025)
by: Yang, An, et al.
Published: (2025)
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
by: Wang, Weizhi, et al.
Published: (2025)
by: Wang, Weizhi, et al.
Published: (2025)
Qwen2.5-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026)
by: Tang, Shengkun, et al.
Published: (2026)
Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Models
by: Xie, Guangyu, et al.
Published: (2025)
by: Xie, Guangyu, et al.
Published: (2025)
Large Language Models Explore by Latent Distilling
by: Zeng, Yuanhao, et al.
Published: (2026)
by: Zeng, Yuanhao, et al.
Published: (2026)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
Evolving Knowledge Distillation for Lightweight Neural Machine Translation
by: Zhang, Xuewen, et al.
Published: (2026)
by: Zhang, Xuewen, et al.
Published: (2026)
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
by: sukhadia, Vrunda N., et al.
Published: (2025)
by: sukhadia, Vrunda N., et al.
Published: (2025)
On-Policy Context Distillation for Language Models
by: Ye, Tianzhu, et al.
Published: (2026)
by: Ye, Tianzhu, et al.
Published: (2026)
D2LLM: Decomposed and Distilled Large Language Models for Semantic Search
by: Liao, Zihan, et al.
Published: (2024)
by: Liao, Zihan, et al.
Published: (2024)
Distilling Token-Trained Models into Byte-Level Models
by: Bao, Zishuo, et al.
Published: (2026)
by: Bao, Zishuo, et al.
Published: (2026)
Knowledge Distillation with Training Wheels
by: Liu, Guanlin, et al.
Published: (2025)
by: Liu, Guanlin, et al.
Published: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
by: Borchers, Conrad, et al.
Published: (2025)
by: Borchers, Conrad, et al.
Published: (2025)
ICH-Qwen: A Large Language Model Towards Chinese Intangible Cultural Heritage
by: Ye, Wenhao, et al.
Published: (2025)
by: Ye, Wenhao, et al.
Published: (2025)
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
by: Shen, Weizhou, et al.
Published: (2025)
by: Shen, Weizhou, et al.
Published: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
by: Huang, Chengyu, et al.
Published: (2024)
by: Huang, Chengyu, et al.
Published: (2024)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
Fine-Tuning Qwen 2.5 3B for Realistic Movie Dialogue Generation
by: Gupta, Kartik
Published: (2025)
by: Gupta, Kartik
Published: (2025)
Knowledge Distillation of Black-Box Large Language Models
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Leveraging Large Language Models for Enhanced NLP Task Performance through Knowledge Distillation and Optimized Training Strategies
by: Huang, Yining, et al.
Published: (2024)
by: Huang, Yining, et al.
Published: (2024)
Quantification of Large Language Model Distillation
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Structured Agent Distillation for Large Language Model
by: Liu, Jun, et al.
Published: (2025)
by: Liu, Jun, et al.
Published: (2025)
Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment
by: Meng, Xiandong, et al.
Published: (2025)
by: Meng, Xiandong, et al.
Published: (2025)
Similar Items
-
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025) -
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025) -
AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
by: Lyu, Yuanjie, et al.
Published: (2026) -
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning
by: Yue, Yuanhao, et al.
Published: (2024) -
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)