WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Zhaojian, Zhang, Xin, Shang, Ning, Huang, Yangyu, Xu, Can, Zhao, Yishujie, Hu, Wenxiang, Yin, Qiufeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914828744916992
author Yu, Zhaojian
Zhang, Xin
Shang, Ning
Huang, Yangyu
Xu, Can
Zhao, Yishujie
Hu, Wenxiang
Yin, Qiufeng
author_facet Yu, Zhaojian
Zhang, Xin
Shang, Ning
Huang, Yangyu
Xu, Can
Zhao, Yishujie
Hu, Wenxiang
Yin, Qiufeng
contents Recent work demonstrates that, after instruction tuning, Code Large Language Models (Code LLMs) can obtain impressive capabilities to address a wide range of code-related tasks. However, current instruction tuning methods for Code LLMs mainly focus on the traditional code generation task, resulting in poor performance in complex multi-task scenarios. In this paper, we concentrate on multiple code-related tasks and present WaveCoder, a series of Code LLMs trained with Widespread And Versatile Enhanced instruction data. To enable the models to tackle complex code-related tasks, we propose a method to stably generate diverse, high-quality instruction data from open source code dataset in multi-task scenarios and obtain CodeSeaXDataset, a dataset comprising 19,915 instruction instances across 4 code-related tasks, which is aimed at improving the generalization ability of Code LLM. Our experiments demonstrate that WaveCoder models significantly outperform other open-source models in terms of the generalization ability across different code-related tasks. Moreover, WaveCoder-Ultra-6.7B presents the state-of-the-art generalization abilities on a wide range of code-related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14187
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
Yu, Zhaojian
Zhang, Xin
Shang, Ning
Huang, Yangyu
Xu, Can
Zhao, Yishujie
Hu, Wenxiang
Yin, Qiufeng
Computation and Language
Artificial Intelligence
Software Engineering
Recent work demonstrates that, after instruction tuning, Code Large Language Models (Code LLMs) can obtain impressive capabilities to address a wide range of code-related tasks. However, current instruction tuning methods for Code LLMs mainly focus on the traditional code generation task, resulting in poor performance in complex multi-task scenarios. In this paper, we concentrate on multiple code-related tasks and present WaveCoder, a series of Code LLMs trained with Widespread And Versatile Enhanced instruction data. To enable the models to tackle complex code-related tasks, we propose a method to stably generate diverse, high-quality instruction data from open source code dataset in multi-task scenarios and obtain CodeSeaXDataset, a dataset comprising 19,915 instruction instances across 4 code-related tasks, which is aimed at improving the generalization ability of Code LLM. Our experiments demonstrate that WaveCoder models significantly outperform other open-source models in terms of the generalization ability across different code-related tasks. Moreover, WaveCoder-Ultra-6.7B presents the state-of-the-art generalization abilities on a wide range of code-related tasks.
title WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2312.14187