Watermarking Should Be Treated as a Monitoring Primitive
Fuente:
arXiv
Saved in:
| Main Authors: | Aremu, Toluwani, Lukas, Nils, Zhang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Safety Monitoring of Language Models via Activation Watermarking
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
by: Aremu, Toluwani, et al.
Published: (2025)
by: Aremu, Toluwani, et al.
Published: (2025)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
by: Yang, Mengmeng, et al.
Published: (2024)
by: Yang, Mengmeng, et al.
Published: (2024)
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)
by: Bagchi, Avi, et al.
Published: (2025)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
by: Bicakci, Kemal
Published: (2026)
by: Bicakci, Kemal
Published: (2026)
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
by: Ly, Reachsak, et al.
Published: (2026)
by: Ly, Reachsak, et al.
Published: (2026)
A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures
by: Wei, Peng, et al.
Published: (2026)
by: Wei, Peng, et al.
Published: (2026)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Adversarial Augmentation and Active Sampling for Robust Cyber Anomaly Detection
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
by: Moreno-Vera, Felipe
Published: (2024)
by: Moreno-Vera, Felipe
Published: (2024)
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
by: Nguyen, Thanh Tam, et al.
Published: (2024)
by: Nguyen, Thanh Tam, et al.
Published: (2024)
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers
by: Burgess, Mark A., et al.
Published: (2024)
by: Burgess, Mark A., et al.
Published: (2024)
The New Frontier of Cybersecurity: Emerging Threats and Innovations
by: Dave, Daksh, et al.
Published: (2023)
by: Dave, Daksh, et al.
Published: (2023)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
SoK: On the Offensive Potential of AI
by: Schröer, Saskia Laura, et al.
Published: (2024)
by: Schröer, Saskia Laura, et al.
Published: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
SecGenAI: Enhancing Security of Cloud-based Generative AI Applications within Australian Critical Technologies of National Interest
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Differentially Private Data Release on Graphs: Inefficiencies and Unfairness
by: Fioretto, Ferdinando, et al.
Published: (2024)
by: Fioretto, Ferdinando, et al.
Published: (2024)
Robustness and Cybersecurity in the EU Artificial Intelligence Act
by: Nolte, Henrik, et al.
Published: (2025)
by: Nolte, Henrik, et al.
Published: (2025)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
by: Shvartzshnaider, Yan, et al.
Published: (2024)
by: Shvartzshnaider, Yan, et al.
Published: (2024)
FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning
by: Sanchez Jr., David, et al.
Published: (2025)
by: Sanchez Jr., David, et al.
Published: (2025)
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
by: Ezzeddine, Fatima, et al.
Published: (2024)
by: Ezzeddine, Fatima, et al.
Published: (2024)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
by: Egea, David, et al.
Published: (2025)
by: Egea, David, et al.
Published: (2025)
Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy
by: Kulynych, Bogdan, et al.
Published: (2025)
by: Kulynych, Bogdan, et al.
Published: (2025)
PUFFLE: Balancing Privacy, Utility, and Fairness in Federated Learning
by: Corbucci, Luca, et al.
Published: (2024)
by: Corbucci, Luca, et al.
Published: (2024)
Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
by: Jandali, Yaman, et al.
Published: (2025)
by: Jandali, Yaman, et al.
Published: (2025)
Downstream Trade-offs of a Family of Text Watermarks
by: Ajith, Anirudh, et al.
Published: (2023)
by: Ajith, Anirudh, et al.
Published: (2023)
Watermarking Makes Language Models Radioactive
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
by: Shao, Shuo, et al.
Published: (2024)
by: Shao, Shuo, et al.
Published: (2024)
Similar Items
-
Robust Safety Monitoring of Language Models via Activation Watermarking
by: Aremu, Toluwani, et al.
Published: (2026) -
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024) -
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
by: Aremu, Toluwani, et al.
Published: (2025) -
Privacy at a Price: Exploring its Dual Impact on AI Fairness
by: Yang, Mengmeng, et al.
Published: (2024) -
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)