Reasoning Structure Matters for Safety Alignment of Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | In, Yeonjun, Kim, Wonjoong, Park, Sangwu, Park, Chanyoung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
by: Kim, Wonjoong, et al.
Published: (2026)
by: Kim, Wonjoong, et al.
Published: (2026)
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
by: In, Yeonjun, et al.
Published: (2026)
by: In, Yeonjun, et al.
Published: (2026)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
by: Kim, Wonjoong, et al.
Published: (2024)
by: Kim, Wonjoong, et al.
Published: (2024)
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
by: Kim, Sein, et al.
Published: (2026)
by: Kim, Sein, et al.
Published: (2026)
Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
by: In, Yeonjun, et al.
Published: (2026)
by: In, Yeonjun, et al.
Published: (2026)
DSLR: Diversity Enhancement and Structure Learning for Rehearsal-based Graph Continual Learning
by: Choi, Seungyoon, et al.
Published: (2024)
by: Choi, Seungyoon, et al.
Published: (2024)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
by: Lee, Seanie, et al.
Published: (2026)
by: Lee, Seanie, et al.
Published: (2026)
Revisiting Fake News Detection: Towards Temporality-aware Evaluation by Leveraging Engagement Earliness
by: Kim, Junghoon, et al.
Published: (2024)
by: Kim, Junghoon, et al.
Published: (2024)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2024)
by: Kim, Kibum, et al.
Published: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
by: Kim, Kibum, et al.
Published: (2023)
by: Kim, Kibum, et al.
Published: (2023)
Structured Debate Improves Corporate Credit Reasoning in Financial AI
by: Lee, Yoonjin, et al.
Published: (2025)
by: Lee, Yoonjin, et al.
Published: (2025)
Reasoning Model is Stubborn: Diagnosing Instruction Overriding in Reasoning Models
by: Jang, Doohyuk, et al.
Published: (2025)
by: Jang, Doohyuk, et al.
Published: (2025)
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
by: Lee, Jeongeun, et al.
Published: (2026)
by: Lee, Jeongeun, et al.
Published: (2026)
Interpretable Prototype-based Graph Information Bottleneck
by: Seo, Sangwoo, et al.
Published: (2023)
by: Seo, Sangwoo, et al.
Published: (2023)
Machine Collective Intelligence for Explainable Scientific Discovery
by: Na, Gyoung S., et al.
Published: (2026)
by: Na, Gyoung S., et al.
Published: (2026)
IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
by: Noh, Heewoong, et al.
Published: (2025)
by: Noh, Heewoong, et al.
Published: (2025)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
by: Kim, Dongyoung, et al.
Published: (2026)
by: Kim, Dongyoung, et al.
Published: (2026)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
by: Hwang, Yeonjun, et al.
Published: (2026)
by: Hwang, Yeonjun, et al.
Published: (2026)
Progressive Multi-Agent Reasoning for Biological Perturbation Prediction
by: Kim, Hyomin, et al.
Published: (2026)
by: Kim, Hyomin, et al.
Published: (2026)
Token-Efficient Item Representation via Images for LLM Recommender Systems
by: Kim, Kibum, et al.
Published: (2025)
by: Kim, Kibum, et al.
Published: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
by: Park, Jueon, et al.
Published: (2026)
by: Park, Jueon, et al.
Published: (2026)
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
by: Son, Yejin, et al.
Published: (2025)
by: Son, Yejin, et al.
Published: (2025)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
by: Seo, Sangwoo, et al.
Published: (2025)
by: Seo, Sangwoo, et al.
Published: (2025)
Vision Language Model is NOT All You Need: Augmentation Strategies for Molecule Language Models
by: Lee, Namkyeong, et al.
Published: (2024)
by: Lee, Namkyeong, et al.
Published: (2024)
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Electron-Informed Coarse-Graining Molecular Representation Learning for Real-World Molecular Physics
by: Na, Gyoung S., et al.
Published: (2026)
by: Na, Gyoung S., et al.
Published: (2026)
Rewarding Structural Conformance of Reasoning using Process Mining
by: Lee, Yongjae, et al.
Published: (2025)
by: Lee, Yongjae, et al.
Published: (2025)
Unsupervised Episode Generation for Graph Meta-learning
by: Jung, Jihyeong, et al.
Published: (2023)
by: Jung, Jihyeong, et al.
Published: (2023)
MobiCLR: Mobility Time Series Contrastive Learning for Urban Region Representations
by: Kim, Namwoo, et al.
Published: (2025)
by: Kim, Namwoo, et al.
Published: (2025)
ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
by: Park, Minbae, et al.
Published: (2025)
by: Park, Minbae, et al.
Published: (2025)
Zero-shot Commonsense Reasoning over Machine Imagination
by: Park, Hyuntae, et al.
Published: (2024)
by: Park, Hyuntae, et al.
Published: (2024)
Similar Items
-
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025) -
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025) -
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
by: Kim, Wonjoong, et al.
Published: (2026) -
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
by: In, Yeonjun, et al.
Published: (2026) -
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
by: Kim, Wonjoong, et al.
Published: (2024)