Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Tong, Du, Yangkai, Ma, Tengfei, Wu, Lingfei, Zhang, Xuhong, Ji, Shouling, Wang, Wenhai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916524214714368
author Ye, Tong
Du, Yangkai
Ma, Tengfei
Wu, Lingfei
Zhang, Xuhong
Ji, Shouling
Wang, Wenhai
author_facet Ye, Tong
Du, Yangkai
Ma, Tengfei
Wu, Lingfei
Zhang, Xuhong
Ji, Shouling
Wang, Wenhai
contents Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ''low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, with AUROC scores increasing by 20.5% on the APPS benchmark and 29.1% on the MBPP benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16133
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
Ye, Tong
Du, Yangkai
Ma, Tengfei
Wu, Lingfei
Zhang, Xuhong
Ji, Shouling
Wang, Wenhai
Software Engineering
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ''low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, with AUROC scores increasing by 20.5% on the APPS benchmark and 29.1% on the MBPP benchmark.
title Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2405.16133