Robust Safety Monitoring of Language Models via Activation Watermarking
Fuente:
arXiv
Saved in:
| Main Authors: | Aremu, Toluwani, Ognev, Daniil, Poppi, Samuele, Lukas, Nils |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Watermarking Should Be Treated as a Monitoring Primitive
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
by: Aremu, Toluwani, et al.
Published: (2025)
by: Aremu, Toluwani, et al.
Published: (2025)
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models
by: Alam, Mohammed Talha, et al.
Published: (2025)
by: Alam, Mohammed Talha, et al.
Published: (2025)
Unforgeable Watermarks for Language Models via Robust Signatures
by: Lin, Huijia, et al.
Published: (2026)
by: Lin, Huijia, et al.
Published: (2026)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)
by: Bagchi, Avi, et al.
Published: (2025)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
Robustness and Cybersecurity in the EU Artificial Intelligence Act
by: Nolte, Henrik, et al.
Published: (2025)
by: Nolte, Henrik, et al.
Published: (2025)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
by: Shvartzshnaider, Yan, et al.
Published: (2024)
by: Shvartzshnaider, Yan, et al.
Published: (2024)
Adversarial Augmentation and Active Sampling for Robust Cyber Anomaly Detection
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
by: Benabderrahmane, Sidahmed, et al.
Published: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
by: Egea, David, et al.
Published: (2025)
by: Egea, David, et al.
Published: (2025)
Watermarking Makes Language Models Radioactive
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
by: Rozenfeld, Shir, et al.
Published: (2026)
by: Rozenfeld, Shir, et al.
Published: (2026)
Watermarking Diffusion Language Models
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Duwak: Dual Watermarks in Large Language Models
by: Zhu, Chaoyi, et al.
Published: (2024)
by: Zhu, Chaoyi, et al.
Published: (2024)
Watermark Stealing in Large Language Models
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
by: Golowich, Noah, et al.
Published: (2024)
by: Golowich, Noah, et al.
Published: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
by: Gu, Tianle, et al.
Published: (2025)
by: Gu, Tianle, et al.
Published: (2025)
A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures
by: Wei, Peng, et al.
Published: (2026)
by: Wei, Peng, et al.
Published: (2026)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
by: Aueawatthanaphisut, Aueaphum
Published: (2025)
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Discovering Spoofing Attempts on Language Model Watermarks
by: Gloaguen, Thibaud, et al.
Published: (2024)
by: Gloaguen, Thibaud, et al.
Published: (2024)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
by: Bicakci, Kemal
Published: (2026)
by: Bicakci, Kemal
Published: (2026)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
by: Nguyen, Thanh Tam, et al.
Published: (2024)
by: Nguyen, Thanh Tam, et al.
Published: (2024)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
by: Ezzeddine, Fatima, et al.
Published: (2024)
by: Ezzeddine, Fatima, et al.
Published: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
by: Block, Adam, et al.
Published: (2025)
by: Block, Adam, et al.
Published: (2025)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
by: Dabiriaghdam, Amirhossein, et al.
Published: (2025)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
by: Wu, Linyu, et al.
Published: (2025)
by: Wu, Linyu, et al.
Published: (2025)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
by: Shao, Shuo, et al.
Published: (2024)
by: Shao, Shuo, et al.
Published: (2024)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
by: Li, Yuexin, et al.
Published: (2026)
by: Li, Yuexin, et al.
Published: (2026)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
by: Ly, Reachsak, et al.
Published: (2026)
by: Ly, Reachsak, et al.
Published: (2026)
Similar Items
-
Watermarking Should Be Treated as a Monitoring Primitive
by: Aremu, Toluwani, et al.
Published: (2026) -
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
by: Aremu, Toluwani, et al.
Published: (2025) -
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024) -
SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models
by: Alam, Mohammed Talha, et al.
Published: (2025) -
Unforgeable Watermarks for Language Models via Robust Signatures
by: Lin, Huijia, et al.
Published: (2026)