Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912277286879232 |
|---|---|
| author | Mei, Jian-Ping Zhang, Weibin Chen, Jie Zhang, Xuyun Zhu, Tiantian |
| author_facet | Mei, Jian-Ping Zhang, Weibin Chen, Jie Zhang, Xuyun Zhu, Tiantian |
| contents | Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_12497 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy Mei, Jian-Ping Zhang, Weibin Chen, Jie Zhang, Xuyun Zhu, Tiantian Cryptography and Security Artificial Intelligence Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings. |
| title | Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2503.12497 |