Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhuoshang, Ren, Yubing, Cao, Yanan, Fang, Fang, Li, Xiaoxue, Guo, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WorldCup Sampling for Multi-bit LLM Watermarking
von: Wang, Yidan, et al.
Veröffentlicht: (2026)
von: Wang, Yidan, et al.
Veröffentlicht: (2026)
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
von: Ganesan, Gokul
Veröffentlicht: (2025)
von: Ganesan, Gokul
Veröffentlicht: (2025)
A Watermark for Black-Box Language Models
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
von: Zhu, He, et al.
Veröffentlicht: (2026)
von: Zhu, He, et al.
Veröffentlicht: (2026)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
Exploring Answer Set Programming for Provenance Graph-Based Cyber Threat Detection: A Novel Approach
von: Li, Fang, et al.
Veröffentlicht: (2025)
von: Li, Fang, et al.
Veröffentlicht: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
Multi-use LLM Watermarking and the False Detection Problem
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering
von: An, Haonan, et al.
Veröffentlicht: (2025)
von: An, Haonan, et al.
Veröffentlicht: (2025)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Black-Box Detection of Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
von: Jawad, Huseein, et al.
Veröffentlicht: (2025)
von: Jawad, Huseein, et al.
Veröffentlicht: (2025)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
von: Yang, Ruozhao, et al.
Veröffentlicht: (2026)
von: Yang, Ruozhao, et al.
Veröffentlicht: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
Black-Box Guardrail Reverse-engineering Attack
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
von: Zhang, Ruisi, et al.
Veröffentlicht: (2023)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2023)
Rethinking Backdoor Detection Evaluation for Language Models
von: Yan, Jun, et al.
Veröffentlicht: (2024)
von: Yan, Jun, et al.
Veröffentlicht: (2024)
SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones
von: Wu, Hengyu, et al.
Veröffentlicht: (2026)
von: Wu, Hengyu, et al.
Veröffentlicht: (2026)
ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples
von: Yang, Yunfei, et al.
Veröffentlicht: (2025)
von: Yang, Yunfei, et al.
Veröffentlicht: (2025)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
von: Fu, Yu, et al.
Veröffentlicht: (2023)
von: Fu, Yu, et al.
Veröffentlicht: (2023)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
von: Dang, Kieu, et al.
Veröffentlicht: (2026)
von: Dang, Kieu, et al.
Veröffentlicht: (2026)
Attacks on Third-Party APIs of Large Language Models
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
BinarySelect to Improve Accessibility of Black-Box Attack Research
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach
von: Li, Xiangfang, et al.
Veröffentlicht: (2025)
von: Li, Xiangfang, et al.
Veröffentlicht: (2025)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
von: Pang, Qi, et al.
Veröffentlicht: (2024)
von: Pang, Qi, et al.
Veröffentlicht: (2024)
Efficient and Universal Watermarking for LLM-Generated Code Detection
von: Li, Boquan, et al.
Veröffentlicht: (2024)
von: Li, Boquan, et al.
Veröffentlicht: (2024)
"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
von: Li, Jiate, et al.
Veröffentlicht: (2026)
von: Li, Jiate, et al.
Veröffentlicht: (2026)
Mark My Words: Analyzing and Evaluating Language Model Watermarks
von: Piet, Julien, et al.
Veröffentlicht: (2023)
von: Piet, Julien, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WorldCup Sampling for Multi-bit LLM Watermarking
von: Wang, Yidan, et al.
Veröffentlicht: (2026) -
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
von: Wang, Yidan, et al.
Veröffentlicht: (2025) -
From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
von: Li, Hao, et al.
Veröffentlicht: (2025) -
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
von: Li, Hao, et al.
Veröffentlicht: (2025) -
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)