Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yixiao, Fang, Binxing, Wang, Rui, Zhou, Yinghai, Liu, Yuan, Li, Mohan, Tian, Zhihong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
by: Pang, Kaiyi, et al.
Published: (2024)
by: Pang, Kaiyi, et al.
Published: (2024)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)
by: Fei, Zekun, et al.
Published: (2024)
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
MEA-Defender: A Robust Watermark against Model Extraction Attack
by: Lv, Peizhuo, et al.
Published: (2024)
by: Lv, Peizhuo, et al.
Published: (2024)
Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models
by: Bai, Minhao, et al.
Published: (2024)
by: Bai, Minhao, et al.
Published: (2024)
QUEEN: Query Unlearning against Model Extraction
by: Chen, Huajie, et al.
Published: (2024)
by: Chen, Huajie, et al.
Published: (2024)
Model Inversion Attack against Federated Unlearning
by: Zhou, Lei, et al.
Published: (2025)
by: Zhou, Lei, et al.
Published: (2025)
A Practical Adversarial Attack against Sequence-based Deep Learning Malware Classifiers
by: Tan, Kai, et al.
Published: (2025)
by: Tan, Kai, et al.
Published: (2025)
Decoupled Alignment for Robust Plug-and-Play Adaptation
by: Luo, Haozheng, et al.
Published: (2024)
by: Luo, Haozheng, et al.
Published: (2024)
GFCL: A GRU-based Federated Continual Learning Framework against Data Poisoning Attacks in IoV
by: Talpur, Anum, et al.
Published: (2022)
by: Talpur, Anum, et al.
Published: (2022)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
by: An, Hyeseon, et al.
Published: (2025)
by: An, Hyeseon, et al.
Published: (2025)
Fault Injection and Safe-Error Attack for Extraction of Embedded Neural Network Models
by: Hector, Kevin, et al.
Published: (2023)
by: Hector, Kevin, et al.
Published: (2023)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Watermark Overwriting Attack on StegaStamp algorithm
by: Serzhenko, I. F., et al.
Published: (2025)
by: Serzhenko, I. F., et al.
Published: (2025)
A High-performance Real-time Container File Monitoring Approach Based on Virtual Machine Introspection
by: Tan, Kai, et al.
Published: (2025)
by: Tan, Kai, et al.
Published: (2025)
LymphNode: A Plug-and-Play Access Control Method for Deep Neural Networks
by: Pei, Hanyu, et al.
Published: (2026)
by: Pei, Hanyu, et al.
Published: (2026)
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
by: Koda, Satoru, et al.
Published: (2024)
by: Koda, Satoru, et al.
Published: (2024)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
by: Wang, Chenrui, et al.
Published: (2025)
by: Wang, Chenrui, et al.
Published: (2025)
CUBA: Controlled Untargeted Backdoor Attack against Deep Neural Networks
by: Wu, Yinghao, et al.
Published: (2025)
by: Wu, Yinghao, et al.
Published: (2025)
MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
by: Cheng, Xueqi, et al.
Published: (2025)
by: Cheng, Xueqi, et al.
Published: (2025)
Probabilistically Robust Watermarking of Neural Networks
by: Pautov, Mikhail, et al.
Published: (2024)
by: Pautov, Mikhail, et al.
Published: (2024)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
by: Fan, Kaisheng, et al.
Published: (2026)
by: Fan, Kaisheng, et al.
Published: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
Watermarking Graph Neural Networks via Explanations for Ownership Protection
by: Downer, Jane, et al.
Published: (2025)
by: Downer, Jane, et al.
Published: (2025)
Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games
by: Kim, Juho, et al.
Published: (2026)
by: Kim, Juho, et al.
Published: (2026)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
by: Sahu, Anubhab, et al.
Published: (2026)
by: Sahu, Anubhab, et al.
Published: (2026)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models
by: Zhang, Shuhao, et al.
Published: (2025)
by: Zhang, Shuhao, et al.
Published: (2025)
DeepNcode: Encoding-Based Protection against Bit-Flip Attacks on Neural Networks
by: Velčický, Patrik, et al.
Published: (2024)
by: Velčický, Patrik, et al.
Published: (2024)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
by: Truong, Vu Tuan, et al.
Published: (2026)
by: Truong, Vu Tuan, et al.
Published: (2026)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
by: Huang, Kaibo, et al.
Published: (2026)
by: Huang, Kaibo, et al.
Published: (2026)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Similar Items
-
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024) -
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
by: Pang, Kaiyi, et al.
Published: (2024) -
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024) -
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
by: Wang, Yidan, et al.
Published: (2025) -
MEA-Defender: A Robust Watermark against Model Extraction Attack
by: Lv, Peizhuo, et al.
Published: (2024)