Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rachmil, Oren, Shapira, Avishag, Betser, Roy, Gershon, Itay, Hofman, Omer, Shabtai, Asaf, Elovici, Yuval, Vainshtein, Roman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
by: Betser, Roy, et al.
Published: (2025)
by: Betser, Roy, et al.
Published: (2025)
Provably Protecting Fine-Tuned LLMs from Training Data Extraction while Preserving Utility
by: Segal, Tom, et al.
Published: (2026)
by: Segal, Tom, et al.
Published: (2026)
FRAME : Comprehensive Risk Assessment Framework for Adversarial Machine Learning Threats
by: Shapira, Avishag, et al.
Published: (2025)
by: Shapira, Avishag, et al.
Published: (2025)
DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs
by: Ganon, Ben, et al.
Published: (2024)
by: Ganon, Ben, et al.
Published: (2024)
Real-World Adversarial Attacks on RF-Based Drone Detectors
by: Gazit, Omer, et al.
Published: (2025)
by: Gazit, Omer, et al.
Published: (2025)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
by: Noah, Amit Finkman, et al.
Published: (2024)
by: Noah, Amit Finkman, et al.
Published: (2024)
LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTI
by: Schwartz, Yuval, et al.
Published: (2024)
by: Schwartz, Yuval, et al.
Published: (2024)
DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
by: Segal, Tom, et al.
Published: (2024)
by: Segal, Tom, et al.
Published: (2024)
Mind the Web: The Security of Web Use Agents
by: Shapira, Avishag, et al.
Published: (2025)
by: Shapira, Avishag, et al.
Published: (2025)
RAPID: Robust APT Detection and Investigation Using Context-Aware Deep Learning
by: Amaru, Yonatan, et al.
Published: (2024)
by: Amaru, Yonatan, et al.
Published: (2024)
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
by: Antebi, Sagiv, et al.
Published: (2025)
by: Antebi, Sagiv, et al.
Published: (2025)
SoK: Cybersecurity Assessment of Humanoid Ecosystem
by: Surve, Priyanka Prakash, et al.
Published: (2025)
by: Surve, Priyanka Prakash, et al.
Published: (2025)
RuleGenie: SIEM Detection Rule Set Optimization
by: Shukla, Akansha, et al.
Published: (2025)
by: Shukla, Akansha, et al.
Published: (2025)
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
by: Baras, Amit, et al.
Published: (2023)
by: Baras, Amit, et al.
Published: (2023)
Addressing Key Challenges of Adversarial Attacks and Defenses in the Tabular Domain: A Methodological Framework for Coherence and Consistency
by: Itzhakev, Yael, et al.
Published: (2024)
by: Itzhakev, Yael, et al.
Published: (2024)
Identifying Memorization of Diffusion Models through $p$-Laplace Analysis: Estimators, Bounds and Applications
by: Brokman, Jonathan, et al.
Published: (2025)
by: Brokman, Jonathan, et al.
Published: (2025)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
Detection of Compromised Functions in a Serverless Cloud Environment
by: Lavi, Danielle, et al.
Published: (2024)
by: Lavi, Danielle, et al.
Published: (2024)
GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection, Localization, Reasoning, and Remediation
by: Malul, Ehud, et al.
Published: (2024)
by: Malul, Ehud, et al.
Published: (2024)
A Privacy Enhancing Technique to Evade Detection by Street Video Cameras Without Using Adversarial Accessories
by: Shams, Jacob, et al.
Published: (2025)
by: Shams, Jacob, et al.
Published: (2025)
From Tool Orchestration to Code Execution: A Study of MCP Design Choices
by: Felendler, Yuval, et al.
Published: (2026)
by: Felendler, Yuval, et al.
Published: (2026)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis
by: Brokman, Jonathan, et al.
Published: (2024)
by: Brokman, Jonathan, et al.
Published: (2024)
SHIELD: APT Detection and Intelligent Explanation Using LLM
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
by: Gandhi, Parth Atulbhai, et al.
Published: (2025)
Rogue Cell: Adversarial Attack and Defense in Untrusted O-RAN Setup Exploiting the Traffic Steering xApp
by: Aizikovich, Eran, et al.
Published: (2025)
by: Aizikovich, Eran, et al.
Published: (2025)
DeSparsify: Adversarial Attack Against Token Sparsification Mechanisms in Vision Transformers
by: Yehezkel, Oryan, et al.
Published: (2024)
by: Yehezkel, Oryan, et al.
Published: (2024)
ImpReSS: Implicit Recommender System for Support Conversations
by: Haller, Omri, et al.
Published: (2025)
by: Haller, Omri, et al.
Published: (2025)
MAPS: A Multilingual Benchmark for Agent Performance and Security
by: Hofman, Omer, et al.
Published: (2025)
by: Hofman, Omer, et al.
Published: (2025)
Whitened CLIP as a Likelihood Surrogate of Images and Captions
by: Betser, Roy, et al.
Published: (2025)
by: Betser, Roy, et al.
Published: (2025)
Peacock: UEFI Firmware Runtime Observability Layer for Detection and Response
by: Gorelik, Hadar Cochavi, et al.
Published: (2026)
by: Gorelik, Hadar Cochavi, et al.
Published: (2026)
Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise
by: Fogel, Ariel, et al.
Published: (2026)
by: Fogel, Ariel, et al.
Published: (2026)
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
Rule-ATT&CK Mapper (RAM): Mapping SIEM Rules to TTPs Using LLMs
by: Wudali, Prasanna N., et al.
Published: (2025)
by: Wudali, Prasanna N., et al.
Published: (2025)
UEFI Memory Forensics: A Framework for UEFI Threat Analysis
by: Segal, Kalanit Suzan, et al.
Published: (2025)
by: Segal, Kalanit Suzan, et al.
Published: (2025)
Towards an End-to-End (E2E) Adversarial Learning and Application in the Physical World
by: Biton, Dudi, et al.
Published: (2025)
by: Biton, Dudi, et al.
Published: (2025)
GPT in Sheep's Clothing: The Risk of Customized GPTs
by: Antebi, Sagiv, et al.
Published: (2024)
by: Antebi, Sagiv, et al.
Published: (2024)
KubeGuard: LLM-Assisted Kubernetes Hardening via Configuration Files and Runtime Logs Analysis
by: Cohen, Omri Sgan, et al.
Published: (2025)
by: Cohen, Omri Sgan, et al.
Published: (2025)
AgenTRIM: Tool Risk Mitigation for Agentic AI
by: Betser, Roy, et al.
Published: (2026)
by: Betser, Roy, et al.
Published: (2026)
Similar Items
-
General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
by: Betser, Roy, et al.
Published: (2025) -
Provably Protecting Fine-Tuned LLMs from Training Data Extraction while Preserving Utility
by: Segal, Tom, et al.
Published: (2026) -
FRAME : Comprehensive Risk Assessment Framework for Adversarial Machine Learning Threats
by: Shapira, Avishag, et al.
Published: (2025) -
DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs
by: Ganon, Ben, et al.
Published: (2024) -
Real-World Adversarial Attacks on RF-Based Drone Detectors
by: Gazit, Omer, et al.
Published: (2025)