LLM Safety From Within: Detecting Harmful Content with Internal Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiao, Difan, Liu, Yilun, Yuan, Ye, Tang, Zhenwei, Du, Linfeng, Wu, Haolun, Anderson, Ashton
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910149213421568
author Jiao, Difan
Liu, Yilun
Yuan, Ye
Tang, Zhenwei
Du, Linfeng
Wu, Haolun
Anderson, Ashton
author_facet Jiao, Difan
Liu, Yilun
Yuan, Ye
Tang, Zhenwei
Du, Linfeng
Wu, Haolun
Anderson, Ashton
contents Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and overlook the rich safety-relevant features distributed across internal layers. We present SIREN, a lightweight guard model that harnesses these internal features. By identifying safety neurons via linear probing and combining them through an adaptive layer-weighted strategy, SIREN builds a harmfulness detector from LLM internals without modifying the underlying model. Our comprehensive evaluation shows that SIREN substantially outperforms state-of-the-art open-source guard models across multiple benchmarks while using 250 times fewer trainable parameters. Moreover, SIREN exhibits superior generalization to unseen benchmarks, naturally enables real-time streaming detection, and significantly improves inference efficiency compared to generative guard models. Overall, our results highlight LLM internal states as a promising foundation for practical, high-performance harmfulness detection.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18519
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM Safety From Within: Detecting Harmful Content with Internal Representations
Jiao, Difan
Liu, Yilun
Yuan, Ye
Tang, Zhenwei
Du, Linfeng
Wu, Haolun
Anderson, Ashton
Artificial Intelligence
Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and overlook the rich safety-relevant features distributed across internal layers. We present SIREN, a lightweight guard model that harnesses these internal features. By identifying safety neurons via linear probing and combining them through an adaptive layer-weighted strategy, SIREN builds a harmfulness detector from LLM internals without modifying the underlying model. Our comprehensive evaluation shows that SIREN substantially outperforms state-of-the-art open-source guard models across multiple benchmarks while using 250 times fewer trainable parameters. Moreover, SIREN exhibits superior generalization to unseen benchmarks, naturally enables real-time streaming detection, and significantly improves inference efficiency compared to generative guard models. Overall, our results highlight LLM internal states as a promising foundation for practical, high-performance harmfulness detection.
title LLM Safety From Within: Detecting Harmful Content with Internal Representations
topic Artificial Intelligence
url https://arxiv.org/abs/2604.18519