Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking
Fuente:
arXiv
Salvato in:
| Autori principali: | You, Ziyang, He, Huilong, Yang, Xiaoke, Lu, Xuxing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense
di: You, Ziyang, et al.
Pubblicazione: (2026)
di: You, Ziyang, et al.
Pubblicazione: (2026)
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
di: You, Ziyang, et al.
Pubblicazione: (2026)
di: You, Ziyang, et al.
Pubblicazione: (2026)
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
di: Shen, Huanming, et al.
Pubblicazione: (2025)
di: Shen, Huanming, et al.
Pubblicazione: (2025)
An Undetectable Watermark for Generative Image Models
di: Gunn, Sam, et al.
Pubblicazione: (2024)
di: Gunn, Sam, et al.
Pubblicazione: (2024)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
di: Lian, Zhuotao, et al.
Pubblicazione: (2025)
di: Lian, Zhuotao, et al.
Pubblicazione: (2025)
Make Split, not Hijack: Preventing Feature-Space Hijacking Attacks in Split Learning
di: Khan, Tanveer, et al.
Pubblicazione: (2024)
di: Khan, Tanveer, et al.
Pubblicazione: (2024)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
di: Wang, Hongtao, et al.
Pubblicazione: (2026)
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
di: Ai, Zhenxin, et al.
Pubblicazione: (2026)
di: Ai, Zhenxin, et al.
Pubblicazione: (2026)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
di: Zhang, Yucheng, et al.
Pubblicazione: (2024)
di: Zhang, Yucheng, et al.
Pubblicazione: (2024)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
di: Wu, Yixin, et al.
Pubblicazione: (2025)
di: Wu, Yixin, et al.
Pubblicazione: (2025)
Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark
di: Fei, Zekun, et al.
Pubblicazione: (2024)
di: Fei, Zekun, et al.
Pubblicazione: (2024)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
di: Noever, David, et al.
Pubblicazione: (2024)
di: Noever, David, et al.
Pubblicazione: (2024)
AgentMark: Utility-Preserving Behavioral Watermarking for Agents
di: Huang, Kaibo, et al.
Pubblicazione: (2026)
di: Huang, Kaibo, et al.
Pubblicazione: (2026)
Character-Level Perturbations Disrupt LLM Watermarks
di: Zhang, Zhaoxi, et al.
Pubblicazione: (2025)
di: Zhang, Zhaoxi, et al.
Pubblicazione: (2025)
Watermark Overwriting Attack on StegaStamp algorithm
di: Serzhenko, I. F., et al.
Pubblicazione: (2025)
di: Serzhenko, I. F., et al.
Pubblicazione: (2025)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
di: Zhang, Qingzhao, et al.
Pubblicazione: (2024)
di: Zhang, Qingzhao, et al.
Pubblicazione: (2024)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
di: Tsmindashvili, Tatia, et al.
Pubblicazione: (2025)
di: Tsmindashvili, Tatia, et al.
Pubblicazione: (2025)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
Optimizing Adaptive Attacks against Watermarks for Language Models
di: Diaa, Abdulrahman, et al.
Pubblicazione: (2024)
di: Diaa, Abdulrahman, et al.
Pubblicazione: (2024)
Sequential Behavioral Watermarking for LLM Agents
di: An, Hyeseon, et al.
Pubblicazione: (2026)
di: An, Hyeseon, et al.
Pubblicazione: (2026)
Attacks and Defenses Against LLM Fingerprinting
di: Kurian, Kevin, et al.
Pubblicazione: (2025)
di: Kurian, Kevin, et al.
Pubblicazione: (2025)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
di: Zhu, Hongyu, et al.
Pubblicazione: (2024)
di: Zhu, Hongyu, et al.
Pubblicazione: (2024)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
LLM Watermark Evasion via Bias Inversion
di: Hwang, Jeongyeon, et al.
Pubblicazione: (2025)
di: Hwang, Jeongyeon, et al.
Pubblicazione: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
di: Zhang, Dongsen, et al.
Pubblicazione: (2025)
Membership Inference Attacks Against Vision-Language Models
di: Hu, Yuke, et al.
Pubblicazione: (2025)
di: Hu, Yuke, et al.
Pubblicazione: (2025)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
di: Xu, Yixiao, et al.
Pubblicazione: (2025)
di: Xu, Yixiao, et al.
Pubblicazione: (2025)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
di: An, Hyeseon, et al.
Pubblicazione: (2025)
di: An, Hyeseon, et al.
Pubblicazione: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
di: Suo, Xuchen
Pubblicazione: (2024)
di: Suo, Xuchen
Pubblicazione: (2024)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
di: Kang, Mintong, et al.
Pubblicazione: (2023)
di: Kang, Mintong, et al.
Pubblicazione: (2023)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks
di: Zhou, Tong, et al.
Pubblicazione: (2025)
di: Zhou, Tong, et al.
Pubblicazione: (2025)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
di: Belkhiter, Yannis, et al.
Pubblicazione: (2026)
di: Belkhiter, Yannis, et al.
Pubblicazione: (2026)
Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection
di: Liu, Qingyu, et al.
Pubblicazione: (2026)
di: Liu, Qingyu, et al.
Pubblicazione: (2026)
CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
di: Xu, Rui, et al.
Pubblicazione: (2025)
di: Xu, Rui, et al.
Pubblicazione: (2025)
Design and Implementation of a Secure RAG-Enhanced AI Chatbot for Smart Tourism Customer Service: Defending Against Prompt Injection Attacks -- A Case Study of Hsinchu, Taiwan
di: Shih, Yu-Kai, et al.
Pubblicazione: (2025)
di: Shih, Yu-Kai, et al.
Pubblicazione: (2025)
Vocabulary Attack to Hijack Large Language Model Applications
di: Levi, Patrick, et al.
Pubblicazione: (2024)
di: Levi, Patrick, et al.
Pubblicazione: (2024)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
di: Tong, Haibo, et al.
Pubblicazione: (2025)
di: Tong, Haibo, et al.
Pubblicazione: (2025)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense
di: You, Ziyang, et al.
Pubblicazione: (2026) -
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
di: You, Ziyang, et al.
Pubblicazione: (2026) -
Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
di: Shen, Huanming, et al.
Pubblicazione: (2025) -
An Undetectable Watermark for Generative Image Models
di: Gunn, Sam, et al.
Pubblicazione: (2024) -
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
di: Lian, Zhuotao, et al.
Pubblicazione: (2025)