Functional Subspace Watermarking for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ding, Zikang, Li, Junhao, Wu, Suling, Yao, Junchi, Liu, Hongbo, Hu, Lijie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912974241792000
author Ding, Zikang
Li, Junhao
Wu, Suling
Yao, Junchi
Liu, Hongbo
Hu, Lijie
author_facet Ding, Zikang
Li, Junhao
Wu, Suling
Yao, Junchi
Liu, Hongbo
Hu, Lijie
contents Model watermarking utilizes internal representations to protect the ownership of large language models (LLMs). However, these features inevitably undergo complex distortions during realistic model modifications such as fine-tuning, quantization, or knowledge distillation, making reliable extraction extremely challenging. Despite extensive research on model-side watermarking, existing methods still lack sufficient robustness against parameter-level perturbations. To address this gap, we propose \texttt{\textbf{Functional Subspace Watermarking (FSW)}}, a framework that anchors ownership signals into a low-dimensional functional backbone. Specifically, we first solve a generalized eigenvalue problem to extract a stable functional subspace for watermark injection, while introducing an adaptive spectral truncation strategy to achieve an optimal balance between robustness and model utility. Furthermore, a vector consistency constraint is incorporated to ensure that watermark injection does not compromise the original semantic performance. Extensive experiments across various LLM architectures and datasets demonstrate that our method achieves superior detection accuracy and statistical verifiability under multiple model attacks, maintaining robustness that outperforms existing state-of-the-art (SOTA) methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18793
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Functional Subspace Watermarking for Large Language Models
Ding, Zikang
Li, Junhao
Wu, Suling
Yao, Junchi
Liu, Hongbo
Hu, Lijie
Cryptography and Security
Artificial Intelligence
Model watermarking utilizes internal representations to protect the ownership of large language models (LLMs). However, these features inevitably undergo complex distortions during realistic model modifications such as fine-tuning, quantization, or knowledge distillation, making reliable extraction extremely challenging. Despite extensive research on model-side watermarking, existing methods still lack sufficient robustness against parameter-level perturbations. To address this gap, we propose \texttt{\textbf{Functional Subspace Watermarking (FSW)}}, a framework that anchors ownership signals into a low-dimensional functional backbone. Specifically, we first solve a generalized eigenvalue problem to extract a stable functional subspace for watermark injection, while introducing an adaptive spectral truncation strategy to achieve an optimal balance between robustness and model utility. Furthermore, a vector consistency constraint is incorporated to ensure that watermark injection does not compromise the original semantic performance. Extensive experiments across various LLM architectures and datasets demonstrate that our method achieves superior detection accuracy and statistical verifiability under multiple model attacks, maintaining robustness that outperforms existing state-of-the-art (SOTA) methods.
title Functional Subspace Watermarking for Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.18793