Is This You, LLM? Recognizing AI-written Programs with Multilingual Code Stylometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gurioli, Andrea, Gabbrielli, Maurizio, Zacchiroli, Stefano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910752633257984
author Gurioli, Andrea
Gabbrielli, Maurizio
Zacchiroli, Stefano
author_facet Gurioli, Andrea
Gabbrielli, Maurizio
Zacchiroli, Stefano
contents With the increasing popularity of LLM-based code completers, like GitHub Copilot, the interest in automatically detecting AI-generated code is also increasing-in particular in contexts where the use of LLMs to program is forbidden by policy due to security, intellectual property, or ethical concerns.We introduce a novel technique for AI code stylometry, i.e., the ability to distinguish code generated by LLMs from code written by humans, based on a transformer-based encoder classifier. Differently from previous work, our classifier is capable of detecting AI-written code across 10 different programming languages with a single machine learning model, maintaining high average accuracy across all languages (84.1% $\pm$ 3.8%).Together with the classifier we also release H-AIRosettaMP, a novel open dataset for AI code stylometry tasks, consisting of 121 247 code snippets in 10 popular programming languages, labeled as either human-written or AI-generated. The experimental pipeline (dataset, training code, resulting models) is the first fully reproducible one for the AI code stylometry task. Most notably our experiments rely only on open LLMs, rather than on proprietary/closed ones like ChatGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Is This You, LLM? Recognizing AI-written Programs with Multilingual Code Stylometry
Gurioli, Andrea
Gabbrielli, Maurizio
Zacchiroli, Stefano
Software Engineering
With the increasing popularity of LLM-based code completers, like GitHub Copilot, the interest in automatically detecting AI-generated code is also increasing-in particular in contexts where the use of LLMs to program is forbidden by policy due to security, intellectual property, or ethical concerns.We introduce a novel technique for AI code stylometry, i.e., the ability to distinguish code generated by LLMs from code written by humans, based on a transformer-based encoder classifier. Differently from previous work, our classifier is capable of detecting AI-written code across 10 different programming languages with a single machine learning model, maintaining high average accuracy across all languages (84.1% $\pm$ 3.8%).Together with the classifier we also release H-AIRosettaMP, a novel open dataset for AI code stylometry tasks, consisting of 121 247 code snippets in 10 popular programming languages, labeled as either human-written or AI-generated. The experimental pipeline (dataset, training code, resulting models) is the first fully reproducible one for the AI code stylometry task. Most notably our experiments rely only on open LLMs, rather than on proprietary/closed ones like ChatGPT.
title Is This You, LLM? Recognizing AI-written Programs with Multilingual Code Stylometry
topic Software Engineering
url https://arxiv.org/abs/2412.14611