PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ai, Zhenxin, He, Haiyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
von: Fei, Zekun, et al.
Veröffentlicht: (2024)
von: Fei, Zekun, et al.
Veröffentlicht: (2024)
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026)
von: You, Ziyang, et al.
Veröffentlicht: (2026)
Modification and Generated-Text Detection: Achieving Dual Detection Capabilities for the Outputs of LLM by Watermark
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
von: Shen, Huanming, et al.
Veröffentlicht: (2025)
PASA: Attack Agnostic Unsupervised Adversarial Detection using Prediction & Attribution Sensitivity Analysis
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
von: Liu, Qingyu, et al.
Veröffentlicht: (2026)
von: Liu, Qingyu, et al.
Veröffentlicht: (2026)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
von: Liu, Zesen, et al.
Veröffentlicht: (2024)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
Invariant-based Robust Weights Watermark for Large Language Models
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
von: Guo, Qingxiao, et al.
Veröffentlicht: (2025)
Character-Level Perturbations Disrupt LLM Watermarks
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaoxi, et al.
Veröffentlicht: (2025)
On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation
von: Omidi, Romina, et al.
Veröffentlicht: (2026)
von: Omidi, Romina, et al.
Veröffentlicht: (2026)
Watermark Overwriting Attack on StegaStamp algorithm
von: Serzhenko, I. F., et al.
Veröffentlicht: (2025)
von: Serzhenko, I. F., et al.
Veröffentlicht: (2025)
Optimizing Adaptive Attacks against Watermarks for Language Models
von: Diaa, Abdulrahman, et al.
Veröffentlicht: (2024)
von: Diaa, Abdulrahman, et al.
Veröffentlicht: (2024)
Sequential Behavioral Watermarking for LLM Agents
von: An, Hyeseon, et al.
Veröffentlicht: (2026)
von: An, Hyeseon, et al.
Veröffentlicht: (2026)
Functional Invariants to Watermark Large Transformers
von: Fernandez, Pierre, et al.
Veröffentlicht: (2023)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2023)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
von: Chen, Tianxin, et al.
Veröffentlicht: (2026)
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2025)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
von: Kim, Jungin, et al.
Veröffentlicht: (2025)
von: Kim, Jungin, et al.
Veröffentlicht: (2025)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
von: An, Hyeseon, et al.
Veröffentlicht: (2025)
von: An, Hyeseon, et al.
Veröffentlicht: (2025)
LLM Watermark Evasion via Bias Inversion
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
von: Zhang, Zhixiang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhixiang, et al.
Veröffentlicht: (2026)
Learning to Watermark LLM-generated Text via Reinforcement Learning
von: Xu, Xiaojun, et al.
Veröffentlicht: (2024)
von: Xu, Xiaojun, et al.
Veröffentlicht: (2024)
RTLMarker: Protecting LLM-Generated RTL Copyright via a Hardware Watermarking Framework
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
von: Liu, Tiantian, et al.
Veröffentlicht: (2024)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
von: Xu, Yixiao, et al.
Veröffentlicht: (2025)
Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors
von: Evennou, Gautier, et al.
Veröffentlicht: (2025)
von: Evennou, Gautier, et al.
Veröffentlicht: (2025)
RobWE: Robust Watermark Embedding for Personalized Federated Learning Model Ownership Protection
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models
von: Müller, Andreas, et al.
Veröffentlicht: (2024)
von: Müller, Andreas, et al.
Veröffentlicht: (2024)
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
The Coding Limits of Robust Watermarking for Generative Models
von: Francati, Danilo, et al.
Veröffentlicht: (2025)
von: Francati, Danilo, et al.
Veröffentlicht: (2025)
CODE ACROSTIC: Robust Watermarking for Code Generation
von: Lin, Li, et al.
Veröffentlicht: (2025)
von: Lin, Li, et al.
Veröffentlicht: (2025)
Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
von: Shan, Shawn, et al.
Veröffentlicht: (2023)
von: Shan, Shawn, et al.
Veröffentlicht: (2023)
MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking
von: Jiang, Ya, et al.
Veröffentlicht: (2026)
von: Jiang, Ya, et al.
Veröffentlicht: (2026)
A Unified Framework for LLM Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Targeted Bit-Flip Attacks on LLM-Based Agents
von: Wang, Jialai, et al.
Veröffentlicht: (2026)
von: Wang, Jialai, et al.
Veröffentlicht: (2026)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
von: Xu, Rui, et al.
Veröffentlicht: (2025)
von: Xu, Rui, et al.
Veröffentlicht: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
von: Suo, Xuchen
Veröffentlicht: (2024)
von: Suo, Xuchen
Veröffentlicht: (2024)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
von: Fei, Zekun, et al.
Veröffentlicht: (2024) -
Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
von: You, Ziyang, et al.
Veröffentlicht: (2026) -
Modification and Generated-Text Detection: Achieving Dual Detection Capabilities for the Outputs of LLM by Watermark
von: Cai, Yuhang, et al.
Veröffentlicht: (2025) -
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
von: Shen, Huanming, et al.
Veröffentlicht: (2025) -
PASA: Attack Agnostic Unsupervised Adversarial Detection using Prediction & Attribution Sensitivity Analysis
von: Bhusal, Dipkamal, et al.
Veröffentlicht: (2024)