Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909549231865856 |
|---|---|
| author | Codefuse Team, Ling : Cai, Wenting Cao, Yuchen Chen, Chaoyu Chen, Chen Chen, Siba Cui, Qing Di, Peng Fang, Junpeng Gong, Zi Guo, Ting He, Zhengyu Huang, Yang Li, Cong Li, Jianguo Li, Zheng Lian, Shijie Liu, BingChang Luo, Songshan Mao, Shuo Shen, Min Wu, Jian Yang, Jiaolong Yang, Wenjie Ye, Tong Yu, Hang Zhang, Wei Zhang, Zhenduo Zhao, Hailin Zheng, Xunjin Zhou, Jun |
| author_facet | Codefuse Team, Ling : Cai, Wenting Cao, Yuchen Chen, Chaoyu Chen, Chen Chen, Siba Cui, Qing Di, Peng Fang, Junpeng Gong, Zi Guo, Ting He, Zhengyu Huang, Yang Li, Cong Li, Jianguo Li, Zheng Lian, Shijie Liu, BingChang Luo, Songshan Mao, Shuo Shen, Min Wu, Jian Yang, Jiaolong Yang, Wenjie Ye, Tong Yu, Hang Zhang, Wei Zhang, Zhenduo Zhao, Hailin Zheng, Xunjin Zhou, Jun |
| contents | Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_17793 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM Codefuse Team, Ling : Cai, Wenting Cao, Yuchen Chen, Chaoyu Chen, Chen Chen, Siba Cui, Qing Di, Peng Fang, Junpeng Gong, Zi Guo, Ting He, Zhengyu Huang, Yang Li, Cong Li, Jianguo Li, Zheng Lian, Shijie Liu, BingChang Luo, Songshan Mao, Shuo Shen, Min Wu, Jian Yang, Jiaolong Yang, Wenjie Ye, Tong Yu, Hang Zhang, Wei Zhang, Zhenduo Zhao, Hailin Zheng, Xunjin Zhou, Jun Machine Learning Artificial Intelligence Computation and Language I.2.7 Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}. |
| title | Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM |
| topic | Machine Learning Artificial Intelligence Computation and Language I.2.7 |
| url | https://arxiv.org/abs/2503.17793 |