CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Rui, Chen, Jiawei, Liu, Weizhi, Yin, Zhaoxia, Kong, Cong, Zhang, Xinpeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910144883851264
author Xu, Rui
Chen, Jiawei
Liu, Weizhi
Yin, Zhaoxia
Kong, Cong
Zhang, Xinpeng
author_facet Xu, Rui
Chen, Jiawei
Liu, Weizhi
Yin, Zhaoxia
Kong, Cong
Zhang, Xinpeng
contents The proliferation of open-source code and large language models (LLMs) for code generation has amplified the risks of unauthorized reuse and intellectual property infringement. Source code watermarking offers a potential solution, yet existing methods typically encode watermarks through identifiers, local code patterns, or limited handcrafted edits, leaving them vulnerable to renaming, refactoring, and adaptive watermark removal. These limitations hinder the joint achievement of robustness, capacity, generalization, and deployment efficiency. We propose CLASP, a Code LLM-Assisted Semantic-Preserving watermarking framework that enables training-free, plug-and-play watermarking for source code. CLASP embeds watermark bits within a fixed space of semantics-preserving transformations, enabling automated watermark insertion with higher capacity while remaining reusable across programming languages and less dependent on brittle lexical features. To recover the watermark, CLASP uses reference-code retrieval and differential comparison to identify transformation traces, avoiding task-specific model training while improving robustness to structural edits and adaptive attacks. Experiments across multiple programming languages show that CLASP consistently outperforms existing baselines in watermark extraction accuracy and robustness, while maintaining code quality under both random removal and adaptive de-watermarking attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11251
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
Xu, Rui
Chen, Jiawei
Liu, Weizhi
Yin, Zhaoxia
Kong, Cong
Zhang, Xinpeng
Cryptography and Security
Artificial Intelligence
Machine Learning
The proliferation of open-source code and large language models (LLMs) for code generation has amplified the risks of unauthorized reuse and intellectual property infringement. Source code watermarking offers a potential solution, yet existing methods typically encode watermarks through identifiers, local code patterns, or limited handcrafted edits, leaving them vulnerable to renaming, refactoring, and adaptive watermark removal. These limitations hinder the joint achievement of robustness, capacity, generalization, and deployment efficiency. We propose CLASP, a Code LLM-Assisted Semantic-Preserving watermarking framework that enables training-free, plug-and-play watermarking for source code. CLASP embeds watermark bits within a fixed space of semantics-preserving transformations, enabling automated watermark insertion with higher capacity while remaining reusable across programming languages and less dependent on brittle lexical features. To recover the watermark, CLASP uses reference-code retrieval and differential comparison to identify transformation traces, avoiding task-specific model training while improving robustness to structural edits and adaptive attacks. Experiments across multiple programming languages show that CLASP consistently outperforms existing baselines in watermark extraction accuracy and robustness, while maintaining code quality under both random removal and adaptive de-watermarking attacks.
title CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.11251