Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Hongyu, Liang, Sichu, Hu, Wentao, Li, Fangqi, Jia, Ju, Wang, Shilin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient and Effective Model Extraction
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024)
Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond
von: Li, Fangqi, et al.
Veröffentlicht: (2024)
von: Li, Fangqi, et al.
Veröffentlicht: (2024)
LLM Watermark Evasion via Bias Inversion
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024)
MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
von: Zhou, Tong, et al.
Veröffentlicht: (2025)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
von: Kang, Mintong, et al.
Veröffentlicht: (2023)
von: Kang, Mintong, et al.
Veröffentlicht: (2023)
Defending Against Beta Poisoning Attacks in Machine Learning Models
von: Gulciftci, Nilufer, et al.
Veröffentlicht: (2025)
von: Gulciftci, Nilufer, et al.
Veröffentlicht: (2025)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
von: Jin, Haotian, et al.
Veröffentlicht: (2025)
von: Jin, Haotian, et al.
Veröffentlicht: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
von: Najt, Elle, et al.
Veröffentlicht: (2026)
von: Najt, Elle, et al.
Veröffentlicht: (2026)
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
von: Xue, Zhiyu, et al.
Veröffentlicht: (2024)
von: Xue, Zhiyu, et al.
Veröffentlicht: (2024)
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
von: Lu, Shilin, et al.
Veröffentlicht: (2024)
von: Lu, Shilin, et al.
Veröffentlicht: (2024)
Quantifying and Defending against Privacy Threats on Federated Knowledge Graph Embedding
von: Hu, Yuke, et al.
Veröffentlicht: (2023)
von: Hu, Yuke, et al.
Veröffentlicht: (2023)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
von: Wu, Daoyuan, et al.
Veröffentlicht: (2024)
von: Wu, Daoyuan, et al.
Veröffentlicht: (2024)
MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
von: Cheng, Xueqi, et al.
Veröffentlicht: (2025)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2025)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
von: Wang, Xunguang, et al.
Veröffentlicht: (2024)
von: Wang, Xunguang, et al.
Veröffentlicht: (2024)
Functional Subspace Watermarking for Large Language Models
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyou, et al.
Veröffentlicht: (2025)
Evading Data Provenance in Deep Neural Networks
von: Zhu, Hongyu, et al.
Veröffentlicht: (2025)
von: Zhu, Hongyu, et al.
Veröffentlicht: (2025)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
von: Lin, Junda, et al.
Veröffentlicht: (2026)
von: Lin, Junda, et al.
Veröffentlicht: (2026)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
Watermarking Techniques for Large Language Models: A Survey
von: Liang, Yuqing, et al.
Veröffentlicht: (2024)
von: Liang, Yuqing, et al.
Veröffentlicht: (2024)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026)
von: You, Ziyang, et al.
Veröffentlicht: (2026)
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
von: Braun, Tobias, et al.
Veröffentlicht: (2026)
von: Braun, Tobias, et al.
Veröffentlicht: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
Invariant-based Robust Weights Watermark for Large Language Models
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
On the Weaknesses of Backdoor-based Model Watermarking: An Information-theoretic Perspective
von: Hu, Aoting, et al.
Veröffentlicht: (2024)
von: Hu, Aoting, et al.
Veröffentlicht: (2024)
Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios
von: Lei, Liangqi, et al.
Veröffentlicht: (2025)
von: Lei, Liangqi, et al.
Veröffentlicht: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
von: Xu, Wenpeng
Veröffentlicht: (2026)
von: Xu, Wenpeng
Veröffentlicht: (2026)
Inevitable Trade-off between Watermark Strength and Speculative Sampling Efficiency for Language Models
von: Hu, Zhengmian, et al.
Veröffentlicht: (2024)
von: Hu, Zhengmian, et al.
Veröffentlicht: (2024)
Multi-Designated Detector Watermarking for Language Models
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
RobWE: Robust Watermark Embedding for Personalized Federated Learning Model Ownership Protection
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Watermarking Visual Concepts for Diffusion Models
von: Lei, Liangqi, et al.
Veröffentlicht: (2024)
von: Lei, Liangqi, et al.
Veröffentlicht: (2024)
AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection
von: Gao, Jianbo, et al.
Veröffentlicht: (2025)
von: Gao, Jianbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient and Effective Model Extraction
von: Zhu, Hongyu, et al.
Veröffentlicht: (2024) -
Revisiting the Information Capacity of Neural Network Watermarks: Upper Bound Estimation and Beyond
von: Li, Fangqi, et al.
Veröffentlicht: (2024) -
LLM Watermark Evasion via Bias Inversion
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025) -
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
von: Zhang, Zhuomeng, et al.
Veröffentlicht: (2024) -
MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks
von: Zhou, Tong, et al.
Veröffentlicht: (2025)