GAVEL: Towards Rule-Based Safety Through Activation Monitoring

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rozenfeld, Shir, Pankajakshan, Rahul, Zloczower, Itay, Lenga, Eyal, Gressel, Gilad, Mirsky, Yisroel
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915969639645184
author Rozenfeld, Shir
Pankajakshan, Rahul
Zloczower, Itay
Lenga, Eyal
Gressel, Gilad
Mirsky, Yisroel
author_facet Rozenfeld, Shir
Pankajakshan, Rahul
Zloczower, Itay
Lenga, Eyal
Gressel, Gilad
Mirsky, Yisroel
contents Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. However, existing activation safety approaches, trained on broad misuse datasets, struggle with poor precision, limited flexibility, and lack of interpretability. This paper introduces a new paradigm: rule-based activation safety, inspired by rule-sharing practices in cybersecurity. We propose modeling activations as cognitive elements (CEs), fine-grained, interpretable factors such as 'making a threat' and 'payment processing', that can be composed to capture nuanced, domain-specific behaviors with higher precision. Building on this representation, we present a practical framework that defines predicate rules over CEs and detects violations in real time. This enables practitioners to configure and update safeguards without retraining models or detectors, while supporting transparency and auditability. Our results show that compositional rule-based activation safety improves precision, supports domain customization, and lays the groundwork for scalable, interpretable, and auditable AI governance. We open source GAVEL and introduce GAVEL Studio, an interactive rule authoring and management tool. Code and datasets are available at github.com/Offensive-AI-Lab/gavel.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19768
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Rozenfeld, Shir
Pankajakshan, Rahul
Zloczower, Itay
Lenga, Eyal
Gressel, Gilad
Mirsky, Yisroel
Artificial Intelligence
Cryptography and Security
Machine Learning
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. However, existing activation safety approaches, trained on broad misuse datasets, struggle with poor precision, limited flexibility, and lack of interpretability. This paper introduces a new paradigm: rule-based activation safety, inspired by rule-sharing practices in cybersecurity. We propose modeling activations as cognitive elements (CEs), fine-grained, interpretable factors such as 'making a threat' and 'payment processing', that can be composed to capture nuanced, domain-specific behaviors with higher precision. Building on this representation, we present a practical framework that defines predicate rules over CEs and detects violations in real time. This enables practitioners to configure and update safeguards without retraining models or detectors, while supporting transparency and auditability. Our results show that compositional rule-based activation safety improves precision, supports domain customization, and lays the groundwork for scalable, interpretable, and auditable AI governance. We open source GAVEL and introduce GAVEL Studio, an interactive rule authoring and management tool. Code and datasets are available at github.com/Offensive-AI-Lab/gavel.
title GAVEL: Towards Rule-Based Safety Through Activation Monitoring
topic Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2601.19768