Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Codefuse, Team, Ling, :, Cai, Wenting, Cao, Yuchen, Chen, Chaoyu, Chen, Chen, Chen, Siba, Cui, Qing, Di, Peng, Fang, Junpeng, Gong, Zi, Guo, Ting, He, Zhengyu, Huang, Yang, Li, Cong, Li, Jianguo, Li, Zheng, Lian, Shijie, Liu, BingChang, Luo, Songshan, Mao, Shuo, Shen, Min, Wu, Jian, Yang, Jiaolong, Yang, Wenjie, Ye, Tong, Yu, Hang, Zhang, Wei, Zhang, Zhenduo, Zhao, Hailin, Zheng, Xunjin, Zhou, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909549231865856
author Codefuse
Team, Ling
:
Cai, Wenting
Cao, Yuchen
Chen, Chaoyu
Chen, Chen
Chen, Siba
Cui, Qing
Di, Peng
Fang, Junpeng
Gong, Zi
Guo, Ting
He, Zhengyu
Huang, Yang
Li, Cong
Li, Jianguo
Li, Zheng
Lian, Shijie
Liu, BingChang
Luo, Songshan
Mao, Shuo
Shen, Min
Wu, Jian
Yang, Jiaolong
Yang, Wenjie
Ye, Tong
Yu, Hang
Zhang, Wei
Zhang, Zhenduo
Zhao, Hailin
Zheng, Xunjin
Zhou, Jun
author_facet Codefuse
Team, Ling
:
Cai, Wenting
Cao, Yuchen
Chen, Chaoyu
Chen, Chen
Chen, Siba
Cui, Qing
Di, Peng
Fang, Junpeng
Gong, Zi
Guo, Ting
He, Zhengyu
Huang, Yang
Li, Cong
Li, Jianguo
Li, Zheng
Lian, Shijie
Liu, BingChang
Luo, Songshan
Mao, Shuo
Shen, Min
Wu, Jian
Yang, Jiaolong
Yang, Wenjie
Ye, Tong
Yu, Hang
Zhang, Wei
Zhang, Zhenduo
Zhao, Hailin
Zheng, Xunjin
Zhou, Jun
contents Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
Codefuse
Team, Ling
:
Cai, Wenting
Cao, Yuchen
Chen, Chaoyu
Chen, Chen
Chen, Siba
Cui, Qing
Di, Peng
Fang, Junpeng
Gong, Zi
Guo, Ting
He, Zhengyu
Huang, Yang
Li, Cong
Li, Jianguo
Li, Zheng
Lian, Shijie
Liu, BingChang
Luo, Songshan
Mao, Shuo
Shen, Min
Wu, Jian
Yang, Jiaolong
Yang, Wenjie
Ye, Tong
Yu, Hang
Zhang, Wei
Zhang, Zhenduo
Zhao, Hailin
Zheng, Xunjin
Zhou, Jun
Machine Learning
Artificial Intelligence
Computation and Language
I.2.7
Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}.
title Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
topic Machine Learning
Artificial Intelligence
Computation and Language
I.2.7
url https://arxiv.org/abs/2503.17793