InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yutong, Huang, Di, Shi, Wenxuan, Wang, Wei, Gao, Lingzhe, Liu, Shihao, Nan, Ziyuan, Yuan, Kaizhao, Zhang, Rui, Zhang, Xishan, Du, Zidong, Guo, Qi, Pu, Yewen, Yin, Dawei, Hu, Xing, Chen, Yunji
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909429359706112
author Wu, Yutong
Huang, Di
Shi, Wenxuan
Wang, Wei
Gao, Lingzhe
Liu, Shihao
Nan, Ziyuan
Yuan, Kaizhao
Zhang, Rui
Zhang, Xishan
Du, Zidong
Guo, Qi
Pu, Yewen
Yin, Dawei
Hu, Xing
Chen, Yunji
author_facet Wu, Yutong
Huang, Di
Shi, Wenxuan
Wang, Wei
Gao, Lingzhe
Liu, Shihao
Nan, Ziyuan
Yuan, Kaizhao
Zhang, Rui
Zhang, Xishan
Du, Zidong
Guo, Qi
Pu, Yewen
Yin, Dawei
Hu, Xing
Chen, Yunji
contents Recent advancements in open-source code large language models (LLMs) have been driven by fine-tuning on the data generated from powerful closed-source LLMs, which are expensive to obtain. This paper explores whether it is possible to use a fine-tuned open-source model to generate additional data to augment its instruction-tuning dataset. We make two observations: (1) A code snippet can serve as the response to different instructions. (2) Instruction-tuned code LLMs perform better at translating code into instructions than the reverse. Based on these observations, we propose Inverse-Instruct, a data augmentation technique that uses a fine-tuned LLM to generate additional instructions of code responses from its own training dataset. The additional instruction-response pairs are added to the original dataset, and a stronger code LLM can be obtained by fine-tuning on the augmented dataset. We empirically validate Inverse-Instruct on a range of open-source code models (e.g. CodeLlama-Python and DeepSeek-Coder) and benchmarks (e.g., HumanEval(+), MBPP(+), DS-1000 and MultiPL-E), showing it consistently improves the base models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05700
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
Wu, Yutong
Huang, Di
Shi, Wenxuan
Wang, Wei
Gao, Lingzhe
Liu, Shihao
Nan, Ziyuan
Yuan, Kaizhao
Zhang, Rui
Zhang, Xishan
Du, Zidong
Guo, Qi
Pu, Yewen
Yin, Dawei
Hu, Xing
Chen, Yunji
Computation and Language
Artificial Intelligence
Software Engineering
Recent advancements in open-source code large language models (LLMs) have been driven by fine-tuning on the data generated from powerful closed-source LLMs, which are expensive to obtain. This paper explores whether it is possible to use a fine-tuned open-source model to generate additional data to augment its instruction-tuning dataset. We make two observations: (1) A code snippet can serve as the response to different instructions. (2) Instruction-tuned code LLMs perform better at translating code into instructions than the reverse. Based on these observations, we propose Inverse-Instruct, a data augmentation technique that uses a fine-tuned LLM to generate additional instructions of code responses from its own training dataset. The additional instruction-response pairs are added to the original dataset, and a stronger code LLM can be obtained by fine-tuning on the augmented dataset. We empirically validate Inverse-Instruct on a range of open-source code models (e.g. CodeLlama-Python and DeepSeek-Coder) and benchmarks (e.g., HumanEval(+), MBPP(+), DS-1000 and MultiPL-E), showing it consistently improves the base models.
title InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
topic Computation and Language
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2407.05700