Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jungin, Park, Shinwoo, Han, Yo-Sub
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912886202302464
author Kim, Jungin
Park, Shinwoo
Han, Yo-Sub
author_facet Kim, Jungin
Park, Shinwoo
Han, Yo-Sub
contents Identifying LLM-generated code through watermarking poses a challenge in preserving functional correctness. Previous methods rely on the assumption that watermarking high-entropy tokens effectively maintains output quality. Our analysis reveals a fundamental limitation of this assumption: syntax-critical tokens such as keywords often exhibit the highest entropy, making existing approaches vulnerable to logic corruption. We present STONE, a syntax-aware watermarking method that embeds watermarks only in non-syntactic tokens and preserves code integrity. For rigorous evaluation, we also introduce STEM, a comprehensive metric that balances three critical dimensions: correctness, detectability, and imperceptibility. Across Python, C++, and Java, STONE preserves correctness, sustains strong detectability, and achieves balanced performance with minimal computational overhead. Our implementation is available at https://github.com/inistory/STONE-watermarking.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
Kim, Jungin
Park, Shinwoo
Han, Yo-Sub
Cryptography and Security
Artificial Intelligence
Identifying LLM-generated code through watermarking poses a challenge in preserving functional correctness. Previous methods rely on the assumption that watermarking high-entropy tokens effectively maintains output quality. Our analysis reveals a fundamental limitation of this assumption: syntax-critical tokens such as keywords often exhibit the highest entropy, making existing approaches vulnerable to logic corruption. We present STONE, a syntax-aware watermarking method that embeds watermarks only in non-syntactic tokens and preserves code integrity. For rigorous evaluation, we also introduce STEM, a comprehensive metric that balances three critical dimensions: correctness, detectability, and imperceptibility. Across Python, C++, and Java, STONE preserves correctness, sustains strong detectability, and achieves balanced performance with minimal computational overhead. Our implementation is available at https://github.com/inistory/STONE-watermarking.
title Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2502.18851