CODE ACROSTIC: Robust Watermarking for Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Li, Xin, Siyuan, Cao, Yang, Cao, Xiaochun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909966208598016
author Lin, Li
Xin, Siyuan
Cao, Yang
Cao, Xiaochun
author_facet Lin, Li
Xin, Siyuan
Cao, Yang
Cao, Xiaochun
contents Watermarking large language models (LLMs) is vital for preventing their misuse, including the fabrication of fake news, plagiarism, and spam. It is especially important to watermark LLM-generated code, as it often contains intellectual property.However, we found that existing methods for watermarking LLM-generated code fail to address comment removal attack.In such cases, an attacker can simply remove the comments from the generated code without affecting its functionality, significantly reducing the effectiveness of current code-watermarking techniques.On the other hand, injecting a watermark into code is challenging because, as previous works have noted, most code represents a low-entropy scenario compared to natural language. Our approach to addressing this issue involves leveraging prior knowledge to distinguish between low-entropy and high-entropy parts of the code, as indicated by a Cue List of words.We then inject the watermark guided by this Cue List, achieving higher detectability and usability than existing methods.We evaluated our proposed method on HumanEvaland compared our method with three state-of-the-art code watermarking techniques. The results demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CODE ACROSTIC: Robust Watermarking for Code Generation
Lin, Li
Xin, Siyuan
Cao, Yang
Cao, Xiaochun
Cryptography and Security
Artificial Intelligence
Watermarking large language models (LLMs) is vital for preventing their misuse, including the fabrication of fake news, plagiarism, and spam. It is especially important to watermark LLM-generated code, as it often contains intellectual property.However, we found that existing methods for watermarking LLM-generated code fail to address comment removal attack.In such cases, an attacker can simply remove the comments from the generated code without affecting its functionality, significantly reducing the effectiveness of current code-watermarking techniques.On the other hand, injecting a watermark into code is challenging because, as previous works have noted, most code represents a low-entropy scenario compared to natural language. Our approach to addressing this issue involves leveraging prior knowledge to distinguish between low-entropy and high-entropy parts of the code, as indicated by a Cue List of words.We then inject the watermark guided by this Cue List, achieving higher detectability and usability than existing methods.We evaluated our proposed method on HumanEvaland compared our method with three state-of-the-art code watermarking techniques. The results demonstrate the effectiveness of our approach.
title CODE ACROSTIC: Robust Watermarking for Code Generation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2512.14753