Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mei, Jian-Ping, Zhang, Weibin, Chen, Jie, Zhang, Xuyun, Zhu, Tiantian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912277286879232
author Mei, Jian-Ping
Zhang, Weibin
Chen, Jie
Zhang, Xuyun
Zhu, Tiantian
author_facet Mei, Jian-Ping
Zhang, Weibin
Chen, Jie
Zhang, Xuyun
Zhu, Tiantian
contents Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy
Mei, Jian-Ping
Zhang, Weibin
Chen, Jie
Zhang, Xuyun
Zhu, Tiantian
Cryptography and Security
Artificial Intelligence
Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings.
title Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2503.12497