A Survey on Model Extraction Attacks and Defenses for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Kaixiang, Li, Lincan, Ding, Kaize, Gong, Neil Zhenqiang, Zhao, Yue, Dong, Yushun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918085849513984
author Zhao, Kaixiang
Li, Lincan
Ding, Kaize
Gong, Neil Zhenqiang
Zhao, Yue
Dong, Yushun
author_facet Zhao, Kaixiang
Li, Lincan
Ding, Kaize
Gong, Neil Zhenqiang
Zhao, Yue
Dong, Yushun
contents Model extraction attacks pose significant security threats to deployed language models, potentially compromising intellectual property and user privacy. This survey provides a comprehensive taxonomy of LLM-specific extraction attacks and defenses, categorizing attacks into functionality extraction, training data extraction, and prompt-targeted attacks. We analyze various attack methodologies including API-based knowledge distillation, direct querying, parameter recovery, and prompt stealing techniques that exploit transformer architectures. We then examine defense mechanisms organized into model protection, data privacy protection, and prompt-targeted strategies, evaluating their effectiveness across different deployment scenarios. We propose specialized metrics for evaluating both attack effectiveness and defense performance, addressing the specific challenges of generative language models. Through our analysis, we identify critical limitations in current approaches and propose promising research directions, including integrated attack methodologies and adaptive defense mechanisms that balance security with model utility. This work serves NLP researchers, ML engineers, and security professionals seeking to protect language models in production environments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Model Extraction Attacks and Defenses for Large Language Models
Zhao, Kaixiang
Li, Lincan
Ding, Kaize
Gong, Neil Zhenqiang
Zhao, Yue
Dong, Yushun
Cryptography and Security
Artificial Intelligence
Machine Learning
Model extraction attacks pose significant security threats to deployed language models, potentially compromising intellectual property and user privacy. This survey provides a comprehensive taxonomy of LLM-specific extraction attacks and defenses, categorizing attacks into functionality extraction, training data extraction, and prompt-targeted attacks. We analyze various attack methodologies including API-based knowledge distillation, direct querying, parameter recovery, and prompt stealing techniques that exploit transformer architectures. We then examine defense mechanisms organized into model protection, data privacy protection, and prompt-targeted strategies, evaluating their effectiveness across different deployment scenarios. We propose specialized metrics for evaluating both attack effectiveness and defense performance, addressing the specific challenges of generative language models. Through our analysis, we identify critical limitations in current approaches and propose promising research directions, including integrated attack methodologies and adaptive defense mechanisms that balance security with model utility. This work serves NLP researchers, ML engineers, and security professionals seeking to protect language models in production environments.
title A Survey on Model Extraction Attacks and Defenses for Large Language Models
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.22521