When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
Fuente:
arXiv
Saved in:
| Main Authors: | Kakkar, Ishita, Zhang, Enze, Uppaal, Rheeya, Hu, Junjie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
by: Uppaal, Rheeya, et al.
Published: (2024)
by: Uppaal, Rheeya, et al.
Published: (2024)
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
by: Uppaal, Rheeya, et al.
Published: (2024)
by: Uppaal, Rheeya, et al.
Published: (2024)
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
by: Uppaal, Rheeya, et al.
Published: (2025)
by: Uppaal, Rheeya, et al.
Published: (2025)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
by: Zhang, Zhaoxin, et al.
Published: (2025)
by: Zhang, Zhaoxin, et al.
Published: (2025)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
by: Rakshit, Sushrita, et al.
Published: (2026)
by: Rakshit, Sushrita, et al.
Published: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
by: Zhu, Dong-Hai, et al.
Published: (2025)
by: Zhu, Dong-Hai, et al.
Published: (2025)
Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education
by: Syamkumar, Anand, et al.
Published: (2024)
by: Syamkumar, Anand, et al.
Published: (2024)
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
by: Tu, Minzhu, et al.
Published: (2026)
by: Tu, Minzhu, et al.
Published: (2026)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)
by: Madhusudhan, Nishanth, et al.
Published: (2026)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
by: Zhao, Xingyu, et al.
Published: (2026)
by: Zhao, Xingyu, et al.
Published: (2026)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
by: Li, Moxin, et al.
Published: (2024)
by: Li, Moxin, et al.
Published: (2024)
Simulated Ignorance Fails: A Systematic Study of LLM Behaviors on Forecasting Problems Before Model Knowledge Cutoff
by: Li, Zehan, et al.
Published: (2026)
by: Li, Zehan, et al.
Published: (2026)
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Diffusion Language Models Know the Answer Before Decoding
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
by: Wang, Changyue, et al.
Published: (2025)
by: Wang, Changyue, et al.
Published: (2025)
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions
by: Zhang, Xiaoyun, et al.
Published: (2024)
by: Zhang, Xiaoyun, et al.
Published: (2024)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
by: Gu, Hexiang, et al.
Published: (2025)
by: Gu, Hexiang, et al.
Published: (2025)
EEE-QA: Exploring Effective and Efficient Question-Answer Representations
by: Hu, Zhanghao, et al.
Published: (2024)
by: Hu, Zhanghao, et al.
Published: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
by: Chen, Yi-Chang, et al.
Published: (2026)
by: Chen, Yi-Chang, et al.
Published: (2026)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
by: Zhang, Shaoqing, et al.
Published: (2024)
by: Zhang, Shaoqing, et al.
Published: (2024)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
by: Zhu, Zihao, et al.
Published: (2025)
by: Zhu, Zihao, et al.
Published: (2025)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
by: Zhan, Weidong, et al.
Published: (2025)
by: Zhan, Weidong, et al.
Published: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
by: Li, Jidong, et al.
Published: (2025)
by: Li, Jidong, et al.
Published: (2025)
Similar Items
-
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
by: Uppaal, Rheeya, et al.
Published: (2024) -
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
by: Uppaal, Rheeya, et al.
Published: (2024) -
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
by: Uppaal, Rheeya, et al.
Published: (2025) -
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025) -
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
by: Zhang, Zhaoxin, et al.
Published: (2025)