Optimizing Adaptive Attacks against Watermarks for Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Diaa, Abdulrahman, Aremu, Toluwani, Lukas, Nils |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Watermarking Should Be Treated as a Monitoring Primitive
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026)
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026)
Robust Safety Monitoring of Language Models via Activation Watermarking
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026)
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026)
Leveraging Optimization for Adaptive Attacks on Image Watermarks
von: Lukas, Nils, et al.
Veröffentlicht: (2023)
von: Lukas, Nils, et al.
Veröffentlicht: (2023)
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
von: Aremu, Toluwani, et al.
Veröffentlicht: (2025)
von: Aremu, Toluwani, et al.
Veröffentlicht: (2025)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
von: Fei, Zekun, et al.
Veröffentlicht: (2024)
von: Fei, Zekun, et al.
Veröffentlicht: (2024)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
von: Wang, Chenrui, et al.
Veröffentlicht: (2025)
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models
von: Zhang, Shuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Shuhao, et al.
Veröffentlicht: (2026)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
Watermarking Language Models for Many Adaptive Users
von: Cohen, Aloni, et al.
Veröffentlicht: (2024)
von: Cohen, Aloni, et al.
Veröffentlicht: (2024)
Watermark Overwriting Attack on StegaStamp algorithm
von: Serzhenko, I. F., et al.
Veröffentlicht: (2025)
von: Serzhenko, I. F., et al.
Veröffentlicht: (2025)
SoK: Robustness in Large Language Models against Jailbreak Attacks
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
Optimized Couplings for Watermarking Large Language Models
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
von: Tsur, Dor, et al.
Veröffentlicht: (2025)
Multi-Designated Detector Watermarking for Language Models
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
von: Huang, Zhengan, et al.
Veröffentlicht: (2024)
Functional Subspace Watermarking for Large Language Models
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
von: Ding, Zikang, et al.
Veröffentlicht: (2026)
Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models
von: Bai, Minhao, et al.
Veröffentlicht: (2024)
von: Bai, Minhao, et al.
Veröffentlicht: (2024)
Model Inversion Attack against Federated Unlearning
von: Zhou, Lei, et al.
Veröffentlicht: (2025)
von: Zhou, Lei, et al.
Veröffentlicht: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
von: Song, Weiming, et al.
Veröffentlicht: (2026)
von: Song, Weiming, et al.
Veröffentlicht: (2026)
Watermarking Techniques for Large Language Models: A Survey
von: Liang, Yuqing, et al.
Veröffentlicht: (2024)
von: Liang, Yuqing, et al.
Veröffentlicht: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
von: Yan, Xiaobei, et al.
Veröffentlicht: (2025)
von: Yan, Xiaobei, et al.
Veröffentlicht: (2025)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
Large Language Model Watermark Stealing With Mixed Integer Programming
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2024)
Vaporizer: Breaking Watermarking Schemes for Large Language Model Outputs
von: Ng, Jonathan Hong Jin, et al.
Veröffentlicht: (2026)
von: Ng, Jonathan Hong Jin, et al.
Veröffentlicht: (2026)
Invariant-based Robust Weights Watermark for Large Language Models
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
von: Li, Yanjie, et al.
Veröffentlicht: (2025)
von: Li, Yanjie, et al.
Veröffentlicht: (2025)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026)
von: You, Ziyang, et al.
Veröffentlicht: (2026)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
von: An, Hyeseon, et al.
Veröffentlicht: (2025)
von: An, Hyeseon, et al.
Veröffentlicht: (2025)
CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models
von: Zhang, Shuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Shuhao, et al.
Veröffentlicht: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
von: Liu, Qingyu, et al.
Veröffentlicht: (2026)
von: Liu, Qingyu, et al.
Veröffentlicht: (2026)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
FedCC: Robust Federated Learning against Model Poisoning Attacks
von: Jeong, Hyejun, et al.
Veröffentlicht: (2022)
von: Jeong, Hyejun, et al.
Veröffentlicht: (2022)
Watermarking Diffusion Language Models
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
StealthInk: A Multi-bit and Stealthy Watermark for Large Language Models
von: Jiang, Ya, et al.
Veröffentlicht: (2025)
von: Jiang, Ya, et al.
Veröffentlicht: (2025)
Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2025)
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2025)
A Survey of Fragile Model Watermarking
von: Gao, Zhenzhe, et al.
Veröffentlicht: (2024)
von: Gao, Zhenzhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Watermarking Should Be Treated as a Monitoring Primitive
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026) -
Robust Safety Monitoring of Language Models via Activation Watermarking
von: Aremu, Toluwani, et al.
Veröffentlicht: (2026) -
Leveraging Optimization for Adaptive Attacks on Image Watermarks
von: Lukas, Nils, et al.
Veröffentlicht: (2023) -
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
von: Aremu, Toluwani, et al.
Veröffentlicht: (2025) -
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)