Black-Box Detection of Language Model Watermarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gloaguen, Thibaud, Jovanović, Nikola, Staab, Robin, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Watermarking Diffusion Language Models
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
Discovering Spoofing Attempts on Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
Towards Watermarking of Open-Source LLMs
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
LLM Fingerprinting via Semantically Conditioned Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
A Unified Framework for LLM Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Watermark Stealing in Large Language Models
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
Ward: Provable RAG Dataset Inference via LLM Watermarks
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
SoK: Data Minimization in Machine Learning
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
Watermarking Autoregressive Image Generation
von: Jovanović, Nikola, et al.
Veröffentlicht: (2025)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2025)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
von: De Muri, Giovanni, et al.
Veröffentlicht: (2025)
von: De Muri, Giovanni, et al.
Veröffentlicht: (2025)
A Watermark for Black-Box Language Models
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
von: Bahri, Dara, et al.
Veröffentlicht: (2024)
Hiding in Plain Sight: Disguising Data Stealing Attacks in Federated Learning
von: Garov, Kostadin, et al.
Veröffentlicht: (2023)
von: Garov, Kostadin, et al.
Veröffentlicht: (2023)
Exploiting LLM Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
Mind the Gap: A Practical Attack on GGUF Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
von: Souček, Tomáš, et al.
Veröffentlicht: (2025)
von: Souček, Tomáš, et al.
Veröffentlicht: (2025)
Black-Box Adversarial Attacks on LLM-Based Code Completion
von: Jenko, Slobodan, et al.
Veröffentlicht: (2024)
von: Jenko, Slobodan, et al.
Veröffentlicht: (2024)
The Challenge of Identifying the Origin of Black-Box Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
Large Language Models for Code: Security Hardening and Adversarial Testing
von: He, Jingxuan, et al.
Veröffentlicht: (2023)
von: He, Jingxuan, et al.
Veröffentlicht: (2023)
DeepEclipse: How to Break White-Box DNN-Watermarking Schemes
von: Pegoraro, Alessandro, et al.
Veröffentlicht: (2024)
von: Pegoraro, Alessandro, et al.
Veröffentlicht: (2024)
Publicly-Detectable Watermarking for Language Models
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2023)
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2023)
Signal Watermark on Large Language Models
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2024)
BaxBench: Can LLMs Generate Correct and Secure Backends?
von: Vero, Mark, et al.
Veröffentlicht: (2025)
von: Vero, Mark, et al.
Veröffentlicht: (2025)
Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection
von: Atiiq, Syafiq Al, et al.
Veröffentlicht: (2026)
von: Atiiq, Syafiq Al, et al.
Veröffentlicht: (2026)
Traceable Black-box Watermarks for Federated Learning
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
Ideal Attribution and Faithful Watermarks for Language Models
von: Song, Min Jae, et al.
Veröffentlicht: (2025)
von: Song, Min Jae, et al.
Veröffentlicht: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
Refined Detection for Gumbel Watermarking
von: Lattimore, Tor
Veröffentlicht: (2026)
von: Lattimore, Tor
Veröffentlicht: (2026)
Breaking Distortion-free Watermarks in Large Language Models
von: Reynolds, Shayleen, et al.
Veröffentlicht: (2025)
von: Reynolds, Shayleen, et al.
Veröffentlicht: (2025)
Large Language Models are Advanced Anonymizers
von: Staab, Robin, et al.
Veröffentlicht: (2024)
von: Staab, Robin, et al.
Veröffentlicht: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Distributional Black-Box Model Inversion Attack with Multi-Agent Reinforcement Learning
von: Bao, Huan, et al.
Veröffentlicht: (2024)
von: Bao, Huan, et al.
Veröffentlicht: (2024)
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
von: Abdulkadir, Perry
Veröffentlicht: (2025)
von: Abdulkadir, Perry
Veröffentlicht: (2025)
Watermarking Language Models through Language Models
von: Dasgupta, Agnibh, et al.
Veröffentlicht: (2024)
von: Dasgupta, Agnibh, et al.
Veröffentlicht: (2024)
On the Learnability of Watermarks for Language Models
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2025)
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2025)
SimKey: A Semantically Aware Key Module for Watermarking Language Models
von: Kodama, Shingo, et al.
Veröffentlicht: (2025)
von: Kodama, Shingo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Watermarking Diffusion Language Models
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025) -
Discovering Spoofing Attempts on Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024) -
Towards Watermarking of Open-Source LLMs
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025) -
LLM Fingerprinting via Semantically Conditioned Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025) -
A Unified Framework for LLM Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)