Magicoder: Empowering Code Generation with OSS-Instruct

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yuxiang, Wang, Zhe, Liu, Jiawei, Ding, Yifeng, Zhang, Lingming
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911909656133632
author Wei, Yuxiang
Wang, Zhe
Liu, Jiawei
Ding, Yifeng
Zhang, Lingming
author_facet Wei, Yuxiang
Wang, Zhe
Liu, Jiawei
Ding, Yifeng
Zhang, Lingming
contents We introduce Magicoder, a series of fully open-source (code, weights, and data) Large Language Models (LLMs) for code that significantly closes the gap with top code models while having no more than 7B parameters. Magicoder models are trained on 75K synthetic instruction data using OSS-Instruct, a novel approach to enlightening LLMs with open-source code snippets to generate diverse instruction data for code. Our main motivation is to mitigate the inherent bias of the synthetic data generated by LLMs through the wealth of open-source references for the production of more realistic and controllable data. The orthogonality of OSS-Instruct and other data generation methods like Evol-Instruct further enables us to build an enhanced MagicoderS. Both Magicoder and MagicoderS substantially outperform state-of-the-art code models with similar or even larger sizes on a wide range of coding benchmarks. Notably, MagicoderS-CL-7B based on CodeLlama even surpasses the prominent ChatGPT on HumanEval+ (66.5 vs. 65.9 in pass@1 ). Overall, OSS-Instruct opens a new direction for crafting diverse synthetic instruction data for code using abundant open-source references.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02120
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Magicoder: Empowering Code Generation with OSS-Instruct
Wei, Yuxiang
Wang, Zhe
Liu, Jiawei
Ding, Yifeng
Zhang, Lingming
Computation and Language
Artificial Intelligence
Software Engineering
We introduce Magicoder, a series of fully open-source (code, weights, and data) Large Language Models (LLMs) for code that significantly closes the gap with top code models while having no more than 7B parameters. Magicoder models are trained on 75K synthetic instruction data using OSS-Instruct, a novel approach to enlightening LLMs with open-source code snippets to generate diverse instruction data for code. Our main motivation is to mitigate the inherent bias of the synthetic data generated by LLMs through the wealth of open-source references for the production of more realistic and controllable data. The orthogonality of OSS-Instruct and other data generation methods like Evol-Instruct further enables us to build an enhanced MagicoderS. Both Magicoder and MagicoderS substantially outperform state-of-the-art code models with similar or even larger sizes on a wide range of coding benchmarks. Notably, MagicoderS-CL-7B based on CodeLlama even surpasses the prominent ChatGPT on HumanEval+ (66.5 vs. 65.9 in pass@1 ). Overall, OSS-Instruct opens a new direction for crafting diverse synthetic instruction data for code using abundant open-source references.
title Magicoder: Empowering Code Generation with OSS-Instruct
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2312.02120