SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhary, Maheep, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024)
by: Lan, Michael, et al.
Published: (2024)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
by: More, Abhishek, et al.
Published: (2025)
by: More, Abhishek, et al.
Published: (2025)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
Certifying Knowledge Comprehension in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
Deception Abilities Emerged in Large Language Models
by: Hagendorff, Thilo
Published: (2023)
by: Hagendorff, Thilo
Published: (2023)
Too Big to Fool: Resisting Deception in Language Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024)
by: Wang, Tony T., et al.
Published: (2024)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
by: Ben-Zion, Ziv, et al.
Published: (2025)
by: Ben-Zion, Ziv, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
by: Berezin, Sergei, et al.
Published: (2025)
by: Berezin, Sergei, et al.
Published: (2025)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
Deception Detection from Linguistic and Physiological Data Streams Using Bimodal Convolutional Neural Networks
by: Li, Panfeng, et al.
Published: (2023)
by: Li, Panfeng, et al.
Published: (2023)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
Similar Items
-
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025) -
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023) -
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)