LLM Safety From Within: Detecting Harmful Content with Internal Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Difan, Liu, Yilun, Yuan, Ye, Tang, Zhenwei, Du, Linfeng, Wu, Haolun, Anderson, Ashton |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification
by: Jiao, Difan, et al.
Published: (2023)
by: Jiao, Difan, et al.
Published: (2023)
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
by: Tang, Zhenwei, et al.
Published: (2025)
by: Tang, Zhenwei, et al.
Published: (2025)
ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
Maia-2: A Unified Model for Human-AI Alignment in Chess
by: Tang, Zhenwei, et al.
Published: (2024)
by: Tang, Zhenwei, et al.
Published: (2024)
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
by: Tang, Zhenwei, et al.
Published: (2025)
by: Tang, Zhenwei, et al.
Published: (2025)
ChessQA: Evaluating Large Language Models for Chess Understanding
by: Wen, Qianfeng, et al.
Published: (2025)
by: Wen, Qianfeng, et al.
Published: (2025)
Level Up: Defining and Exploiting Transitional Problems for Curriculum Learning
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
MINER: Mining Multimodal Internal Representation for Efficient Retrieval
by: Li, Weien, et al.
Published: (2026)
by: Li, Weien, et al.
Published: (2026)
Grounded Chess Reasoning in Language Models via Master Distillation
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media
by: Lu, Rui, et al.
Published: (2025)
by: Lu, Rui, et al.
Published: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Beyond Accuracy: An Explainability-Driven Analysis of Harmful Content Detection
by: Dhara, Trishita, et al.
Published: (2026)
by: Dhara, Trishita, et al.
Published: (2026)
TEARS: Textual Representations for Scrutable Recommendations
by: Penaloza, Emiliano, et al.
Published: (2024)
by: Penaloza, Emiliano, et al.
Published: (2024)
From Internal Representations to Text Quality: A Geometric Approach to LLM Evaluation
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
Scaling Tasks, Not Samples: Mastering Humanoid Control through Multi-Task Model-Based Reinforcement Learning
by: Liu, Shaohuai, et al.
Published: (2026)
by: Liu, Shaohuai, et al.
Published: (2026)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
by: Xiao, Wenjie, et al.
Published: (2026)
by: Xiao, Wenjie, et al.
Published: (2026)
More of the Same: Persistent Representational Harms Under Increased Representation
by: Mickel, Jennifer, et al.
Published: (2025)
by: Mickel, Jennifer, et al.
Published: (2025)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
by: Li, Yanhui, et al.
Published: (2025)
by: Li, Yanhui, et al.
Published: (2025)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
by: Zhang, Shaoqing, et al.
Published: (2024)
by: Zhang, Shaoqing, et al.
Published: (2024)
Why Do Large Language Models Generate Harmful Content?
by: Ganguli, Rajesh, et al.
Published: (2026)
by: Ganguli, Rajesh, et al.
Published: (2026)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
by: Stepanov, Ihor, et al.
Published: (2026)
by: Stepanov, Ihor, et al.
Published: (2026)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
by: Zhao, Haochen, et al.
Published: (2024)
by: Zhao, Haochen, et al.
Published: (2024)
Harmful Visual Content Manipulation Matters in Misinformation Detection Under Multimedia Scenarios
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Beyond Message Passing: A Semantic View of Agent Communication Protocols
by: Yuan, Dun, et al.
Published: (2026)
by: Yuan, Dun, et al.
Published: (2026)
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
by: Zheng, Han, et al.
Published: (2026)
by: Zheng, Han, et al.
Published: (2026)
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
by: Wen, Ming, et al.
Published: (2026)
by: Wen, Ming, et al.
Published: (2026)
Internal Representations as Indicators of Hallucinations in Agent Tool Selection
by: Healy, Kait, et al.
Published: (2026)
by: Healy, Kait, et al.
Published: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
by: Asif, Sadia, et al.
Published: (2025)
by: Asif, Sadia, et al.
Published: (2025)
Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts
by: Bo, Jessica Y., et al.
Published: (2026)
by: Bo, Jessica Y., et al.
Published: (2026)
Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
by: Han, Yujin, et al.
Published: (2025)
by: Han, Yujin, et al.
Published: (2025)
Similar Items
-
SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification
by: Jiao, Difan, et al.
Published: (2023) -
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
by: Tang, Zhenwei, et al.
Published: (2025) -
ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
by: Jiao, Difan, et al.
Published: (2026) -
Maia-2: A Unified Model for Human-AI Alignment in Chess
by: Tang, Zhenwei, et al.
Published: (2024) -
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
by: Tang, Zhenwei, et al.
Published: (2025)