Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | An, Li, Liu, Yujian, Liu, Yepeng, Zhang, Yang, Bu, Yuheng, Chang, Shiyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Reinforcement Learning Framework for Robust and Secure LLM Watermarking
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
by: Liu, Yepeng, et al.
Published: (2025)
by: Liu, Yepeng, et al.
Published: (2025)
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
by: Liu, Yepeng, et al.
Published: (2025)
by: Liu, Yepeng, et al.
Published: (2025)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
Membership Inference Attacks Against In-Context Learning
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approach
by: He, Haiyun, et al.
Published: (2024)
by: He, Haiyun, et al.
Published: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
by: Huang, Hanbo, et al.
Published: (2026)
by: Huang, Hanbo, et al.
Published: (2026)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
SPML: A DSL for Defending Language Models Against Prompt Attacks
by: Sharma, Reshabh K, et al.
Published: (2024)
by: Sharma, Reshabh K, et al.
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024)
by: Sha, Zeyang, et al.
Published: (2024)
Distributional Information Embedding: A Framework for Multi-bit Watermarking
by: He, Haiyun, et al.
Published: (2025)
by: He, Haiyun, et al.
Published: (2025)
MPMA: Preference Manipulation Attack Against Model Context Protocol
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature
by: Zhou, Tong, et al.
Published: (2024)
by: Zhou, Tong, et al.
Published: (2024)
Enhance Robustness of Language Models Against Variation Attack through Graph Integration
by: Xiong, Zi, et al.
Published: (2024)
by: Xiong, Zi, et al.
Published: (2024)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
WorldCup Sampling for Multi-bit LLM Watermarking
by: Wang, Yidan, et al.
Published: (2026)
by: Wang, Yidan, et al.
Published: (2026)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
Image Watermarks are Removable Using Controllable Regeneration from Clean Noise
by: Liu, Yepeng, et al.
Published: (2024)
by: Liu, Yepeng, et al.
Published: (2024)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
by: Ganesan, Gokul
Published: (2025)
by: Ganesan, Gokul
Published: (2025)
Securing Genomic Data Against Inference Attacks in Federated Learning Environments
by: Pathade, Chetan, et al.
Published: (2025)
by: Pathade, Chetan, et al.
Published: (2025)
Multi-use LLM Watermarking and the False Detection Problem
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
by: Zhang, Xiaozhe, et al.
Published: (2026)
by: Zhang, Xiaozhe, et al.
Published: (2026)
LLM Anonymization Against Agentic Re-Identification
by: Li, Ziwen, et al.
Published: (2026)
by: Li, Ziwen, et al.
Published: (2026)
RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
by: Huang, Hanbo, et al.
Published: (2025)
by: Huang, Hanbo, et al.
Published: (2025)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
by: Dang, Kieu, et al.
Published: (2026)
by: Dang, Kieu, et al.
Published: (2026)
Similar Items
-
A Reinforcement Learning Framework for Robust and Secure LLM Watermarking
by: An, Li, et al.
Published: (2025) -
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
by: Liu, Yepeng, et al.
Published: (2025) -
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
by: Liu, Yepeng, et al.
Published: (2025) -
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
by: Li, Yuanfan, et al.
Published: (2025) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)