GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Fuente:
arXiv
Saved in:
| Main Authors: | Rozenfeld, Shir, Pankajakshan, Rahul, Zloczower, Itay, Lenga, Eyal, Gressel, Gilad, Mirsky, Yisroel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
by: Gressel, Gilad, et al.
Published: (2025)
by: Gressel, Gilad, et al.
Published: (2025)
Who Owns This Agent? Tracing AI Agents Back to Their Owners
by: Chocron, Ruben, et al.
Published: (2026)
by: Chocron, Ruben, et al.
Published: (2026)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
Mapping LLM Security Landscapes: A Comprehensive Stakeholder Risk Assessment Proposal
by: Pankajakshan, Rahul, et al.
Published: (2024)
by: Pankajakshan, Rahul, et al.
Published: (2024)
PEAS: A Strategy for Crafting Transferable Adversarial Examples
by: Avraham, Bar, et al.
Published: (2024)
by: Avraham, Bar, et al.
Published: (2024)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
by: Hacmon, Yaniv, et al.
Published: (2025)
by: Hacmon, Yaniv, et al.
Published: (2025)
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
by: Bernstein, Shir, et al.
Published: (2025)
by: Bernstein, Shir, et al.
Published: (2025)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
by: Weiss, Roy, et al.
Published: (2024)
by: Weiss, Roy, et al.
Published: (2024)
Efficient Model Extraction via Boundary Sampling
by: Dor, Maor Biton, et al.
Published: (2024)
by: Dor, Maor Biton, et al.
Published: (2024)
ASRJam: Human-Friendly AI Speech Jamming to Prevent Automated Phone Scams
by: Grabovski, Freddie, et al.
Published: (2025)
by: Grabovski, Freddie, et al.
Published: (2025)
Transpose Attack: Stealing Datasets with Bidirectional Training
by: Amit, Guy, et al.
Published: (2023)
by: Amit, Guy, et al.
Published: (2023)
Transferability Ranking of Adversarial Examples
by: Levy, Mosh, et al.
Published: (2022)
by: Levy, Mosh, et al.
Published: (2022)
Robust Safety Monitoring of Language Models via Activation Watermarking
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models
by: Hadad, Eyal, et al.
Published: (2026)
by: Hadad, Eyal, et al.
Published: (2026)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Differentially Private Iterative Screening Rules for Linear Regression
by: Khanna, Amol, et al.
Published: (2025)
by: Khanna, Amol, et al.
Published: (2025)
Uncertainty-Aware Hardware Trojan Detection Using Multimodal Deep Learning
by: Vishwakarma, Rahul, et al.
Published: (2024)
by: Vishwakarma, Rahul, et al.
Published: (2024)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
Smooth Sensitivity for Learning Differentially-Private yet Accurate Rule Lists
by: Ly, Timothée, et al.
Published: (2024)
by: Ly, Timothée, et al.
Published: (2024)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
by: Shuster, Shaun, et al.
Published: (2025)
by: Shuster, Shaun, et al.
Published: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Cutting Through Privacy: A Hyperplane-Based Data Reconstruction Attack in Federated Learning
by: Diana, Francesco, et al.
Published: (2025)
by: Diana, Francesco, et al.
Published: (2025)
Eclectic Rule Extraction for Explainability of Deep Neural Network based Intrusion Detection Systems
by: Ables, Jesse, et al.
Published: (2024)
by: Ables, Jesse, et al.
Published: (2024)
Saffron-1: Safety Inference Scaling
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Towards Family-Grouped Hierarchical Federated Learning on Sub-5KB Models: A Feasibility Study of Privacy-Preserving ECG Monitoring for Ultra-Resource-Constrained Wearables
by: Wu, Hangyu
Published: (2026)
by: Wu, Hangyu
Published: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
In-Context Representation Hijacking
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Privacy-Preserving Data Sharing in Agriculture: Enforcing Policy Rules for Secure and Confidential Data Synthesis
by: Kotal, Anantaa, et al.
Published: (2023)
by: Kotal, Anantaa, et al.
Published: (2023)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
Why Safety Probes Catch Liars But Miss Fanatics
by: Haralambiev, Kristiyan
Published: (2026)
by: Haralambiev, Kristiyan
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Similar Items
-
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026) -
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
by: Gressel, Gilad, et al.
Published: (2025) -
Who Owns This Agent? Tracing AI Agents Back to Their Owners
by: Chocron, Ruben, et al.
Published: (2026) -
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024) -
Mapping LLM Security Landscapes: A Comprehensive Stakeholder Risk Assessment Proposal
by: Pankajakshan, Rahul, et al.
Published: (2024)