Code Comparison Tuning for Code Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Yufan, He, Qiaozhi, Zhuang, Xiaomin, Wu, Zhihua
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929374201118720
author Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
author_facet Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
contents We present Code Comparison Tuning (CCT), a simple and effective tuning method for code large language models (Code LLMs) to better handle subtle code errors. Specifically, we integrate the concept of comparison into instruction tuning, both at the token and sequence levels, enabling the model to discern even the slightest deviations in code. To compare the original code with an erroneous version containing manually added code errors, we use token-level preference loss for detailed token-level comparisons. Additionally, we combine code segments to create a new instruction tuning sample for sequence-level comparisons, enhancing the model's bug-fixing capability. Experimental results on the HumanEvalFix benchmark show that CCT surpasses instruction tuning in pass@1 scores by up to 4 points across diverse code LLMs, and extensive analysis demonstrates the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19121
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Code Comparison Tuning for Code Large Language Models
Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
Computation and Language
We present Code Comparison Tuning (CCT), a simple and effective tuning method for code large language models (Code LLMs) to better handle subtle code errors. Specifically, we integrate the concept of comparison into instruction tuning, both at the token and sequence levels, enabling the model to discern even the slightest deviations in code. To compare the original code with an erroneous version containing manually added code errors, we use token-level preference loss for detailed token-level comparisons. Additionally, we combine code segments to create a new instruction tuning sample for sequence-level comparisons, enhancing the model's bug-fixing capability. Experimental results on the HumanEvalFix benchmark show that CCT surpasses instruction tuning in pass@1 scores by up to 4 points across diverse code LLMs, and extensive analysis demonstrates the effectiveness of our method.
title Code Comparison Tuning for Code Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2403.19121