RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Hanbo, Zhang, Yiran, Zheng, Hao, Gong, Xuan, Li, Yihan, Liu, Lin, Liu, Zhuotao, Liang, Shiyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
by: Huang, Hanbo, et al.
Published: (2026)
by: Huang, Hanbo, et al.
Published: (2026)
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
by: Huang, Hanbo, et al.
Published: (2024)
by: Huang, Hanbo, et al.
Published: (2024)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
Analyzing and Evaluating Unbiased Language Model Watermark
by: Wu, Yihan, et al.
Published: (2025)
by: Wu, Yihan, et al.
Published: (2025)
A Reinforcement Learning Framework for Robust and Secure LLM Watermarking
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
by: Pang, Kaiyi, et al.
Published: (2024)
by: Pang, Kaiyi, et al.
Published: (2024)
An Ensemble Framework for Unbiased Language Model Watermarking
by: Wu, Yihan, et al.
Published: (2025)
by: Wu, Yihan, et al.
Published: (2025)
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
by: Yin, Yu, et al.
Published: (2026)
by: Yin, Yu, et al.
Published: (2026)
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
by: Zhang, Zhixiang, et al.
Published: (2026)
by: Zhang, Zhixiang, et al.
Published: (2026)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
by: Zhang, Bokang, et al.
Published: (2025)
by: Zhang, Bokang, et al.
Published: (2025)
Leveraging Optimization for Adaptive Attacks on Image Watermarks
by: Lukas, Nils, et al.
Published: (2023)
by: Lukas, Nils, et al.
Published: (2023)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
Cluster-Aware Attacks on Graph Watermarks
by: Nemecek, Alexander, et al.
Published: (2025)
by: Nemecek, Alexander, et al.
Published: (2025)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)
by: Liang, Jiacheng, et al.
Published: (2024)
AMUSE: Adaptive Multi-Segment Encoding for Dataset Watermarking
by: Alvar, Saeed Ranjbar, et al.
Published: (2024)
by: Alvar, Saeed Ranjbar, et al.
Published: (2024)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
by: Lian, Zhuotao, et al.
Published: (2025)
by: Lian, Zhuotao, et al.
Published: (2025)
More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles
by: Chen, Ruibo, et al.
Published: (2026)
by: Chen, Ruibo, et al.
Published: (2026)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
by: Wang, Yihan, et al.
Published: (2025)
by: Wang, Yihan, et al.
Published: (2025)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026)
by: Liu, Shi, et al.
Published: (2026)
Distortion-free Watermarks are not Truly Distortion-free under Watermark Key Collisions
by: Wu, Yihan, et al.
Published: (2024)
by: Wu, Yihan, et al.
Published: (2024)
Agentic Privacy-Preserving Machine Learning
by: Zhang, Mengyu, et al.
Published: (2025)
by: Zhang, Mengyu, et al.
Published: (2025)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
A Transfer Attack to Image Watermarks
by: Hu, Yuepeng, et al.
Published: (2024)
by: Hu, Yuepeng, et al.
Published: (2024)
Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack
by: Hui, Siqi, et al.
Published: (2025)
by: Hui, Siqi, et al.
Published: (2025)
FedMUA: Exploring the Vulnerabilities of Federated Learning to Malicious Unlearning Attacks
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
Evaluating the Vulnerability Landscape of LLM-Generated Smart Contracts
by: Do, Hoang Long, et al.
Published: (2026)
by: Do, Hoang Long, et al.
Published: (2026)
MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages
by: Cui, Xuehao, et al.
Published: (2026)
by: Cui, Xuehao, et al.
Published: (2026)
LLM-BSCVM: An LLM-Based Blockchain Smart Contract Vulnerability Management Framework
by: Jin, Yanli, et al.
Published: (2025)
by: Jin, Yanli, et al.
Published: (2025)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)
by: Fei, Zekun, et al.
Published: (2024)
Composability in Watermarking Schemes
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
by: Hastuti, Rochana Prih, et al.
Published: (2025)
by: Hastuti, Rochana Prih, et al.
Published: (2025)
martFL: Enabling Utility-Driven Data Marketplace with a Robust and Verifiable Federated Learning Architecture
by: Li, Qi, et al.
Published: (2023)
by: Li, Qi, et al.
Published: (2023)
Demystifying RCE Vulnerabilities in LLM-Integrated Apps
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
by: Qi, Senmao, et al.
Published: (2025)
by: Qi, Senmao, et al.
Published: (2025)
VideoMark: A Distortion-Free Robust Watermarking Framework for Video Diffusion Models
by: Hu, Xuming, et al.
Published: (2025)
by: Hu, Xuming, et al.
Published: (2025)
Similar Items
-
RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
by: Huang, Hanbo, et al.
Published: (2026) -
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
by: Huang, Hanbo, et al.
Published: (2024) -
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025) -
Analyzing and Evaluating Unbiased Language Model Watermark
by: Wu, Yihan, et al.
Published: (2025) -
A Reinforcement Learning Framework for Robust and Secure LLM Watermarking
by: An, Li, et al.
Published: (2025)