WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
Fuente:
arXiv
Guardado en:
| Autores principales: | Giboulot, Eva, Furon, Teddy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Guidance Watermarking for Diffusion Models
por: Gesny, Enoal, et al.
Publicado: (2025)
por: Gesny, Enoal, et al.
Publicado: (2025)
Proving membership in LLM pretraining data via data watermarks
por: Wei, Johnny Tian-Zheng, et al.
Publicado: (2024)
por: Wei, Johnny Tian-Zheng, et al.
Publicado: (2024)
Watermarking Makes Language Models Radioactive
por: Sander, Tom, et al.
Publicado: (2024)
por: Sander, Tom, et al.
Publicado: (2024)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
por: Zhao, Xuandong, et al.
Publicado: (2024)
por: Zhao, Xuandong, et al.
Publicado: (2024)
Functional Invariants to Watermark Large Transformers
por: Fernandez, Pierre, et al.
Publicado: (2023)
por: Fernandez, Pierre, et al.
Publicado: (2023)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
por: Pang, Qi, et al.
Publicado: (2024)
por: Pang, Qi, et al.
Publicado: (2024)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
por: Zhang, Yuehan, et al.
Publicado: (2024)
por: Zhang, Yuehan, et al.
Publicado: (2024)
Backdoor Attacks on Deep Learning Face Detection
por: Roux, Quentin Le, et al.
Publicado: (2025)
por: Roux, Quentin Le, et al.
Publicado: (2025)
Task-Agnostic Attacks Against Vision Foundation Models
por: Pulfer, Brian, et al.
Publicado: (2025)
por: Pulfer, Brian, et al.
Publicado: (2025)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
por: Lee, Seanie, et al.
Publicado: (2024)
por: Lee, Seanie, et al.
Publicado: (2024)
SoK: On the Survivability of Backdoor Attacks on Unconstrained Face Recognition Systems
por: Roux, Quentin Le, et al.
Publicado: (2025)
por: Roux, Quentin Le, et al.
Publicado: (2025)
LLM Unlearning Should Be Form-Independent
por: Ye, Xiaotian, et al.
Publicado: (2025)
por: Ye, Xiaotian, et al.
Publicado: (2025)
GCG Attack On A Diffusion LLM
por: Neyroud, Ruben, et al.
Publicado: (2025)
por: Neyroud, Ruben, et al.
Publicado: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
por: Goyal, Shubh, et al.
Publicado: (2024)
por: Goyal, Shubh, et al.
Publicado: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
por: Assogba, Yannick, et al.
Publicado: (2026)
por: Assogba, Yannick, et al.
Publicado: (2026)
Localizing Malicious Outputs from CodeLLM
por: Borana, Mayukh, et al.
Publicado: (2025)
por: Borana, Mayukh, et al.
Publicado: (2025)
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
por: Gesny, Enoal, et al.
Publicado: (2026)
por: Gesny, Enoal, et al.
Publicado: (2026)
Secure Seed-Based Multi-bit Watermarking for Diffusion Models from First Principles
por: Gesny, Enoal, et al.
Publicado: (2026)
por: Gesny, Enoal, et al.
Publicado: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
por: Wang, Tianchun, et al.
Publicado: (2024)
por: Wang, Tianchun, et al.
Publicado: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
por: Zhao, Xuandong, et al.
Publicado: (2025)
por: Zhao, Xuandong, et al.
Publicado: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
por: Liu, Ruixuan, et al.
Publicado: (2026)
por: Liu, Ruixuan, et al.
Publicado: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
por: Krishna, Arjun, et al.
Publicado: (2025)
por: Krishna, Arjun, et al.
Publicado: (2025)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
por: Duan, Haohua, et al.
Publicado: (2025)
por: Duan, Haohua, et al.
Publicado: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
por: Zhang, Boyang, et al.
Publicado: (2026)
por: Zhang, Boyang, et al.
Publicado: (2026)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
por: Shafee, Samaneh, et al.
Publicado: (2024)
por: Shafee, Samaneh, et al.
Publicado: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
por: Liang, Jiacheng, et al.
Publicado: (2024)
por: Liang, Jiacheng, et al.
Publicado: (2024)
LLM Dataset Inference: Did you train on my dataset?
por: Maini, Pratyush, et al.
Publicado: (2024)
por: Maini, Pratyush, et al.
Publicado: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
por: Liu, Hongfu, et al.
Publicado: (2024)
por: Liu, Hongfu, et al.
Publicado: (2024)
Robust LLM safeguarding via refusal feature adversarial training
por: Yu, Lei, et al.
Publicado: (2024)
por: Yu, Lei, et al.
Publicado: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
por: Meeus, Matthieu, et al.
Publicado: (2025)
por: Meeus, Matthieu, et al.
Publicado: (2025)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
por: Sander, Tom, et al.
Publicado: (2026)
por: Sander, Tom, et al.
Publicado: (2026)
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
por: Yang, Xikang, et al.
Publicado: (2024)
por: Yang, Xikang, et al.
Publicado: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
por: Chen, Zhiyu, et al.
Publicado: (2024)
por: Chen, Zhiyu, et al.
Publicado: (2024)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
por: Hilel, Almog, et al.
Publicado: (2025)
por: Hilel, Almog, et al.
Publicado: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
por: Cai, Will, et al.
Publicado: (2025)
por: Cai, Will, et al.
Publicado: (2025)
FreqMark: Frequency-Based Watermark for Sentence-Level Detection of LLM-Generated Text
por: Xu, Zhenyu, et al.
Publicado: (2024)
por: Xu, Zhenyu, et al.
Publicado: (2024)
Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
por: Rao, Zixin, et al.
Publicado: (2025)
por: Rao, Zixin, et al.
Publicado: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
por: Wang, Fei, et al.
Publicado: (2025)
por: Wang, Fei, et al.
Publicado: (2025)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
por: Pathmanathan, Pankayaraj, et al.
Publicado: (2024)
Ejemplares similares
-
Guidance Watermarking for Diffusion Models
por: Gesny, Enoal, et al.
Publicado: (2025) -
Proving membership in LLM pretraining data via data watermarks
por: Wei, Johnny Tian-Zheng, et al.
Publicado: (2024) -
Watermarking Makes Language Models Radioactive
por: Sander, Tom, et al.
Publicado: (2024) -
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
por: Zhao, Xuandong, et al.
Publicado: (2024) -
Functional Invariants to Watermark Large Transformers
por: Fernandez, Pierre, et al.
Publicado: (2023)