MEA-Defender: A Robust Watermark against Model Extraction Attack
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lv, Peizhuo, Ma, Hualong, Chen, Kai, Zhou, Jiachen, Zhang, Shengzhi, Liang, Ruigang, Zhu, Shenchen, Li, Pan, Zhang, Yingjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-supervised Learning
von: Lv, Peizhuo, et al.
Veröffentlicht: (2022)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2022)
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
A Model Stealing Attack Against Multi-Exit Networks
von: Pan, Li, et al.
Veröffentlicht: (2023)
von: Pan, Li, et al.
Veröffentlicht: (2023)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
von: Pang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Pang, Kaiyi, et al.
Veröffentlicht: (2024)
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer
von: Zhao, Yue, et al.
Veröffentlicht: (2026)
von: Zhao, Yue, et al.
Veröffentlicht: (2026)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
von: Cheng, Xueqi, et al.
Veröffentlicht: (2025)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
von: Hu, Runyi, et al.
Veröffentlicht: (2024)
Defending Against Neural Network Model Inversion Attacks via Data Poisoning
von: Zhou, Shuai, et al.
Veröffentlicht: (2024)
von: Zhou, Shuai, et al.
Veröffentlicht: (2024)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
von: Yang, Guang, et al.
Veröffentlicht: (2024)
von: Yang, Guang, et al.
Veröffentlicht: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Collective Certified Robustness against Graph Injection Attacks
von: Lai, Yuni, et al.
Veröffentlicht: (2024)
von: Lai, Yuni, et al.
Veröffentlicht: (2024)
ShapeMark: Robust and Diversity-Preserving Watermarking for Diffusion Models
von: Qian, Yuqi, et al.
Veröffentlicht: (2026)
von: Qian, Yuqi, et al.
Veröffentlicht: (2026)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
von: Koda, Satoru, et al.
Veröffentlicht: (2024)
von: Koda, Satoru, et al.
Veröffentlicht: (2024)
Dormant: Defending against Pose-driven Human Image Animation
von: Zhou, Jiachen, et al.
Veröffentlicht: (2024)
von: Zhou, Jiachen, et al.
Veröffentlicht: (2024)
A Defensive Framework Against Adversarial Attacks on Machine Learning-Based Network Intrusion Detection Systems
von: Tafreshian, Benyamin, et al.
Veröffentlicht: (2025)
von: Tafreshian, Benyamin, et al.
Veröffentlicht: (2025)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
Optimizing Adaptive Attacks against Watermarks for Language Models
von: Diaa, Abdulrahman, et al.
Veröffentlicht: (2024)
von: Diaa, Abdulrahman, et al.
Veröffentlicht: (2024)
Defending against Data Poisoning Attacks in Federated Learning via User Elimination
von: Galanis, Nick
Veröffentlicht: (2024)
von: Galanis, Nick
Veröffentlicht: (2024)
Defending against Backdoor Attack on Deep Neural Networks
von: Cheng, Hao, et al.
Veröffentlicht: (2020)
von: Cheng, Hao, et al.
Veröffentlicht: (2020)
Invariant Aggregator for Defending against Federated Backdoor Attacks
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2022)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2022)
SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
von: Yao, Lingfeng, et al.
Veröffentlicht: (2025)
von: Yao, Lingfeng, et al.
Veröffentlicht: (2025)
Backdoor Attacks against Hybrid Classical-Quantum Neural Networks
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
Revisiting Gradient Pruning: A Dual Realization for Defending against Gradient Attacks
von: Xue, Lulu, et al.
Veröffentlicht: (2024)
von: Xue, Lulu, et al.
Veröffentlicht: (2024)
Adversarial Attacks on Reinforcement Learning-based Medical Questionnaire Systems: Input-level Perturbation Strategies and Medical Constraint Validation
von: Liu, Peizhuo
Veröffentlicht: (2025)
von: Liu, Peizhuo
Veröffentlicht: (2025)
Your Agent Can Defend Itself against Backdoor Attacks
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
A Robust Semantics-based Watermark for Large Language Model against Paraphrasing
von: Ren, Jie, et al.
Veröffentlicht: (2023)
von: Ren, Jie, et al.
Veröffentlicht: (2023)
An Automated Attack Investigation Approach Leveraging Threat-Knowledge-Augmented Large Language Models
von: Dai, Rujie, et al.
Veröffentlicht: (2025)
von: Dai, Rujie, et al.
Veröffentlicht: (2025)
PhantomFetch: Obfuscating Loads against Prefetcher Side-Channel Attacks
von: Zhang, Xingzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Xingzhi, et al.
Veröffentlicht: (2025)
A Game Between the Defender and the Attacker for Trigger-based Black-box Model Watermarking
von: Huang, Chaoyue, et al.
Veröffentlicht: (2025)
von: Huang, Chaoyue, et al.
Veröffentlicht: (2025)
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
von: Li, Shuai, et al.
Veröffentlicht: (2023)
von: Li, Shuai, et al.
Veröffentlicht: (2023)
Secure Goal-Oriented Communication: Defending against Eavesdropping Timing Attacks
von: Mason, Federico, et al.
Veröffentlicht: (2025)
von: Mason, Federico, et al.
Veröffentlicht: (2025)
Invariant-based Robust Weights Watermark for Large Language Models
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-supervised Learning
von: Lv, Peizhuo, et al.
Veröffentlicht: (2022) -
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025) -
A Model Stealing Attack Against Multi-Exit Networks
von: Pan, Li, et al.
Veröffentlicht: (2023) -
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
von: Pang, Kaiyi, et al.
Veröffentlicht: (2024) -
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)