How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Su-Hyeon, Jin, Hyundong, Lee, Yejin, Han, Yo-Sub |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
by: Kim, Su-Hyeon, et al.
Published: (2026)
by: Kim, Su-Hyeon, et al.
Published: (2026)
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
by: Jin, Hyundong, et al.
Published: (2026)
by: Jin, Hyundong, et al.
Published: (2026)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
by: Jin, Hyundong, et al.
Published: (2026)
by: Jin, Hyundong, et al.
Published: (2026)
Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
by: Kim, Su-Hyeon, et al.
Published: (2026)
by: Kim, Su-Hyeon, et al.
Published: (2026)
Steering Language Models Before They Speak: Logit-Level Interventions
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Detection of LLM-Paraphrased Code and Identification of the Responsible LLM Using Coding Style Features
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
STAB: Specification-driven Testing for Algorithmic Bottlenecks
by: Lim, Soohan, et al.
Published: (2026)
by: Lim, Soohan, et al.
Published: (2026)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
by: Kim, Su-Hyeon, et al.
Published: (2025)
by: Kim, Su-Hyeon, et al.
Published: (2025)
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025)
by: Kim, Jungin, et al.
Published: (2025)
Sequential Behavioral Watermarking for LLM Agents
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
by: Lee, Yejin, et al.
Published: (2026)
by: Lee, Yejin, et al.
Published: (2026)
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
by: Tang, Peiyuan, et al.
Published: (2025)
by: Tang, Peiyuan, et al.
Published: (2025)
URECA: The Chain of Two Minimum Set Cover Problems exists behind Adaptation to Shifts in Semantic Code Search
by: Choi, Seok-Ung, et al.
Published: (2025)
by: Choi, Seok-Ung, et al.
Published: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Linguistics-Aware Non-Distortionary LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
by: An, Hyeseon, et al.
Published: (2025)
by: An, Hyeseon, et al.
Published: (2025)
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
by: Son, Yejin, et al.
Published: (2025)
by: Son, Yejin, et al.
Published: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
Repairing Regex Vulnerabilities via Localization-Guided Instructions
by: Sung, Sicheol, et al.
Published: (2025)
by: Sung, Sicheol, et al.
Published: (2025)
Continual Learning for Multiple Modalities
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming
by: Sung, Sicheol, et al.
Published: (2025)
by: Sung, Sicheol, et al.
Published: (2025)
Position: AI Safety Must Embrace an Antifragile Perspective
by: Jin, Ming, et al.
Published: (2025)
by: Jin, Ming, et al.
Published: (2025)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
by: Puerto, Haritz, et al.
Published: (2026)
by: Puerto, Haritz, et al.
Published: (2026)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy
by: Li, Zeju, et al.
Published: (2025)
by: Li, Zeju, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Group Pattern Selection Optimization: Let LRMs Pick the Right Pattern for Reasoning
by: Wang, Hanbin, et al.
Published: (2026)
by: Wang, Hanbin, et al.
Published: (2026)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
Similar Items
-
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
by: Kim, Su-Hyeon, et al.
Published: (2026) -
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
by: Jin, Hyundong, et al.
Published: (2026) -
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
by: Lee, Yejin, et al.
Published: (2025) -
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
by: Jin, Hyundong, et al.
Published: (2026) -
Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
by: Kim, Su-Hyeon, et al.
Published: (2026)