$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Kieu, Lai, Phung, Phan, NhatHai, Shen, Yelong, Jin, Ruoming, Khreishah, Abdallah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914114255716352
author Dang, Kieu
Lai, Phung
Phan, NhatHai
Shen, Yelong
Jin, Ruoming
Khreishah, Abdallah
author_facet Dang, Kieu
Lai, Phung
Phan, NhatHai
Shen, Yelong
Jin, Ruoming
Khreishah, Abdallah
contents Large language models (LLMs) demonstrate remarkable capabilities across various tasks. However, their deployment introduces significant risks related to intellectual property. In this context, we focus on model stealing attacks, where adversaries replicate the behaviors of these models to steal services. These attacks are highly relevant to proprietary LLMs and pose serious threats to revenue and financial stability. To mitigate these risks, the watermarking solution embeds imperceptible patterns in LLM outputs, enabling model traceability and intellectual property verification. In this paper, we study the vulnerability of LLM service providers by introducing $δ$-STEAL, a novel model stealing attack that bypasses the service provider's watermark detectors while preserving the adversary's model utility. $δ$-STEAL injects noise into the token embeddings of the adversary's model during fine-tuning in a way that satisfies local differential privacy (LDP) guarantees. The adversary queries the service provider's model to collect outputs and form input-output training pairs. By applying LDP-preserving noise to these pairs, $δ$-STEAL obfuscates watermark signals, making it difficult for the service provider to determine whether its outputs were used, thereby preventing claims of model theft. Our experiments show that $δ$-STEAL with lightweight modifications achieves attack success rates of up to $96.95\%$ without significantly compromising the adversary's model utility. The noise scale in LDP controls the trade-off between attack effectiveness and model utility. This poses a significant risk, as even robust watermarks can be bypassed, allowing adversaries to deceive watermark detectors and undermine current intellectual property protection methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21946
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle $δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
Dang, Kieu
Lai, Phung
Phan, NhatHai
Shen, Yelong
Jin, Ruoming
Khreishah, Abdallah
Cryptography and Security
68T07, 68T50
I.2.6; I.2.7; K.6.5
Large language models (LLMs) demonstrate remarkable capabilities across various tasks. However, their deployment introduces significant risks related to intellectual property. In this context, we focus on model stealing attacks, where adversaries replicate the behaviors of these models to steal services. These attacks are highly relevant to proprietary LLMs and pose serious threats to revenue and financial stability. To mitigate these risks, the watermarking solution embeds imperceptible patterns in LLM outputs, enabling model traceability and intellectual property verification. In this paper, we study the vulnerability of LLM service providers by introducing $δ$-STEAL, a novel model stealing attack that bypasses the service provider's watermark detectors while preserving the adversary's model utility. $δ$-STEAL injects noise into the token embeddings of the adversary's model during fine-tuning in a way that satisfies local differential privacy (LDP) guarantees. The adversary queries the service provider's model to collect outputs and form input-output training pairs. By applying LDP-preserving noise to these pairs, $δ$-STEAL obfuscates watermark signals, making it difficult for the service provider to determine whether its outputs were used, thereby preventing claims of model theft. Our experiments show that $δ$-STEAL with lightweight modifications achieves attack success rates of up to $96.95\%$ without significantly compromising the adversary's model utility. The noise scale in LDP controls the trade-off between attack effectiveness and model utility. This poses a significant risk, as even robust watermarks can be bypassed, allowing adversaries to deceive watermark detectors and undermine current intellectual property protection methods.
title $δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
topic Cryptography and Security
68T07, 68T50
I.2.6; I.2.7; K.6.5
url https://arxiv.org/abs/2510.21946