LLM Watermark Evasion via Bias Inversion
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Jeongyeon, Park, Sangdon, Ok, Jungseul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedBN: Robust Test-Time Adaptation against Malicious Test Samples
by: Park, Hyejin, et al.
Published: (2024)
by: Park, Hyejin, et al.
Published: (2024)
AI Kill Switch for malicious web-based LLM agent
by: Lee, Sechan, et al.
Published: (2025)
by: Lee, Sechan, et al.
Published: (2025)
Making Models Unmergeable via Scaling-Sensitive Loss Landscape
by: Jang, Minwoo, et al.
Published: (2026)
by: Jang, Minwoo, et al.
Published: (2026)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
by: Zhu, Hongyu, et al.
Published: (2024)
by: Zhu, Hongyu, et al.
Published: (2024)
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025)
by: Kim, Jungin, et al.
Published: (2025)
NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
by: An, Hyeseon, et al.
Published: (2025)
by: An, Hyeseon, et al.
Published: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
by: Zhang, Zhaoxi, et al.
Published: (2025)
by: Zhang, Zhaoxi, et al.
Published: (2025)
RTLMarker: Protecting LLM-Generated RTL Copyright via a Hardware Watermarking Framework
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
by: Sun, Luze, et al.
Published: (2026)
by: Sun, Luze, et al.
Published: (2026)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
by: Shen, Huanming, et al.
Published: (2025)
by: Shen, Huanming, et al.
Published: (2025)
MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking
by: Jiang, Ya, et al.
Published: (2026)
by: Jiang, Ya, et al.
Published: (2026)
Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
by: Wang, Chenrui, et al.
Published: (2025)
by: Wang, Chenrui, et al.
Published: (2025)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
by: You, Ziyang, et al.
Published: (2026)
by: You, Ziyang, et al.
Published: (2026)
On Protecting Agentic Systems' Intellectual Property via Watermarking
by: Wang, Liwen, et al.
Published: (2026)
by: Wang, Liwen, et al.
Published: (2026)
Learning to Watermark LLM-generated Text via Reinforcement Learning
by: Xu, Xiaojun, et al.
Published: (2024)
by: Xu, Xiaojun, et al.
Published: (2024)
Ward: Provable RAG Dataset Inference via LLM Watermarks
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
by: Domico, Kyle, et al.
Published: (2025)
by: Domico, Kyle, et al.
Published: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
Modification and Generated-Text Detection: Achieving Dual Detection Capabilities for the Outputs of LLM by Watermark
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation
by: Omidi, Romina, et al.
Published: (2026)
by: Omidi, Romina, et al.
Published: (2026)
Watermarking Graph Neural Networks via Explanations for Ownership Protection
by: Downer, Jane, et al.
Published: (2025)
by: Downer, Jane, et al.
Published: (2025)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
by: Kang, Mintong, et al.
Published: (2023)
by: Kang, Mintong, et al.
Published: (2023)
A Unified Framework for LLM Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
by: Dong, Tian, et al.
Published: (2025)
by: Dong, Tian, et al.
Published: (2025)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
by: Fei, Zekun, et al.
Published: (2024)
by: Fei, Zekun, et al.
Published: (2024)
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
by: Gao, Jiaxin, et al.
Published: (2025)
by: Gao, Jiaxin, et al.
Published: (2025)
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
by: Ai, Zhenxin, et al.
Published: (2026)
by: Ai, Zhenxin, et al.
Published: (2026)
OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization
by: Xing, Jiazheng, et al.
Published: (2025)
by: Xing, Jiazheng, et al.
Published: (2025)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
by: Liu, Qingyu, et al.
Published: (2026)
by: Liu, Qingyu, et al.
Published: (2026)
Hide&Seek: Remove Image Watermarks with Negligible Cost via Pixel-wise Reconstruction
by: Chen, Huajie, et al.
Published: (2026)
by: Chen, Huajie, et al.
Published: (2026)
Detecting Benchmark Contamination Through Watermarking
by: Sander, Tom, et al.
Published: (2025)
by: Sander, Tom, et al.
Published: (2025)
Probabilistically Robust Watermarking of Neural Networks
by: Pautov, Mikhail, et al.
Published: (2024)
by: Pautov, Mikhail, et al.
Published: (2024)
A Survey of Fragile Model Watermarking
by: Gao, Zhenzhe, et al.
Published: (2024)
by: Gao, Zhenzhe, et al.
Published: (2024)
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
by: Gu, Chenxi, et al.
Published: (2026)
by: Gu, Chenxi, et al.
Published: (2026)
EGAN: Evolutional GAN for Ransomware Evasion
by: Commey, Daniel, et al.
Published: (2024)
by: Commey, Daniel, et al.
Published: (2024)
Similar Items
-
MedBN: Robust Test-Time Adaptation against Malicious Test Samples
by: Park, Hyejin, et al.
Published: (2024) -
AI Kill Switch for malicious web-based LLM agent
by: Lee, Sechan, et al.
Published: (2025) -
Making Models Unmergeable via Scaling-Sensitive Loss Landscape
by: Jang, Minwoo, et al.
Published: (2026) -
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024) -
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
by: Zhu, Hongyu, et al.
Published: (2024)