Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
Fuente:
arXiv
Saved in:
| Main Authors: | Kadali, Sri Durga Sai Sowmya, Papalexakis, Evangelos E. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026)
CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection
by: Li, Jerry, et al.
Published: (2025)
by: Li, Jerry, et al.
Published: (2025)
GPT-generated Text Detection: Benchmark Dataset and Tensor-based Detection Method
by: Qazi, Zubair, et al.
Published: (2024)
by: Qazi, Zubair, et al.
Published: (2024)
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)
by: Preiß, Martin
Published: (2025)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method
by: Kilaru, Sai Hemanth, et al.
Published: (2026)
by: Kilaru, Sai Hemanth, et al.
Published: (2026)
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023)
by: Rao, Abhinav, et al.
Published: (2023)
The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
by: Hu, Luoming, et al.
Published: (2026)
by: Hu, Luoming, et al.
Published: (2026)
Clue-Instruct: Text-Based Clue Generation for Educational Crossword Puzzles
by: Zugarini, Andrea, et al.
Published: (2024)
by: Zugarini, Andrea, et al.
Published: (2024)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
by: Patel, Het, et al.
Published: (2025)
by: Patel, Het, et al.
Published: (2025)
ExpertGenQA: Open-ended QA generation in Specialized Domains
by: Shahgir, Haz Sameen, et al.
Published: (2025)
by: Shahgir, Haz Sameen, et al.
Published: (2025)
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
IndicGEC: Powerful Models, or a Measurement Mirage?
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
The Problem with Safety Classification is not just the Models
by: Vajjala, Sowmya
Published: (2025)
by: Vajjala, Sowmya
Published: (2025)
INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Semantic Volume: Quantifying and Detecting both External and Internal Uncertainty in LLMs
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
by: Yang, Zhipeng, et al.
Published: (2025)
by: Yang, Zhipeng, et al.
Published: (2025)
Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
by: Nguyen, Tri, et al.
Published: (2025)
by: Nguyen, Tri, et al.
Published: (2025)
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
by: Shiromani, Shikhar, et al.
Published: (2026)
by: Shiromani, Shikhar, et al.
Published: (2026)
Mitigating Jailbreaks with Intent-Aware LLMs
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Jailbreaking LLMs with Arabic Transliteration and Arabizi
by: Ghanim, Mansour Al, et al.
Published: (2024)
by: Ghanim, Mansour Al, et al.
Published: (2024)
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
by: Nguyen, Tuan T., et al.
Published: (2025)
by: Nguyen, Tuan T., et al.
Published: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026)
by: Patel, Het, et al.
Published: (2026)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025)
by: Kritz, Jeremy, et al.
Published: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
Discrete Stochastic Localization for Non-autoregressive Generation
by: Wu, Yunshu, et al.
Published: (2026)
by: Wu, Yunshu, et al.
Published: (2026)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
by: Meng, Wenlong, et al.
Published: (2025)
by: Meng, Wenlong, et al.
Published: (2025)
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
by: Zhou, Yao, et al.
Published: (2026)
by: Zhou, Yao, et al.
Published: (2026)
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
by: Wang, Zijun, et al.
Published: (2024)
by: Wang, Zijun, et al.
Published: (2024)
Intention Analysis Makes LLMs A Good Jailbreak Defender
by: Zhang, Yuqi, et al.
Published: (2024)
by: Zhang, Yuqi, et al.
Published: (2024)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
Playing Language Game with LLMs Leads to Jailbreaking
by: Peng, Yu, et al.
Published: (2024)
by: Peng, Yu, et al.
Published: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
Text Classification in the LLM Era -- Where do we stand?
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
Similar Items
-
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2026) -
CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025) -
Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection
by: Li, Jerry, et al.
Published: (2025) -
GPT-generated Text Detection: Benchmark Dataset and Tensor-based Detection Method
by: Qazi, Zubair, et al.
Published: (2024) -
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)