Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Chentao, Xu, Xiaojun, Han, Bo, Li, Hang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
Reasoning as an Adaptive Defense for Safety
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Can LLM Safety Be Ensured by Constraining Parameter Regions?
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
von: Zuo, Qian, et al.
Veröffentlicht: (2025)
von: Zuo, Qian, et al.
Veröffentlicht: (2025)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2024)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2025)
Adversarial Reasoning at Jailbreaking Time
von: Sabbaghi, Mahdi, et al.
Veröffentlicht: (2025)
von: Sabbaghi, Mahdi, et al.
Veröffentlicht: (2025)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
von: Zhang, Haonan, et al.
Veröffentlicht: (2026)
von: Zhang, Haonan, et al.
Veröffentlicht: (2026)
Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners
von: Pandey, Rohan, et al.
Veröffentlicht: (2026)
von: Pandey, Rohan, et al.
Veröffentlicht: (2026)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
Plantain: Plan-Answer Interleaved Reasoning
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
Incentivizing LLMs to Self-Verify Their Answers
von: Zhang, Fuxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Fuxiang, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
Multilingual Safety Alignment via Self-Distillation
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
von: Garcia-Gasulla, Dario, et al.
Veröffentlicht: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment
von: Li, Cheryl, et al.
Veröffentlicht: (2025)
von: Li, Cheryl, et al.
Veröffentlicht: (2025)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Curriculum Learning for Safety Alignment
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2026)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
LocalGCL: Local-aware Contrastive Learning for Graphs
von: Jiang, Haojun, et al.
Veröffentlicht: (2024)
von: Jiang, Haojun, et al.
Veröffentlicht: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Course-Correction: Safety Alignment Using Synthetic Preferences
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
von: Kiyani, Shayan, et al.
Veröffentlicht: (2026)
von: Kiyani, Shayan, et al.
Veröffentlicht: (2026)
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
von: Ma, Jiachen, et al.
Veröffentlicht: (2026)
von: Ma, Jiachen, et al.
Veröffentlicht: (2026)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
von: Wang, Chenyang, et al.
Veröffentlicht: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Gradients as an Action: Towards Communication-Efficient Federated Recommender Systems via Adaptive Action Sharing
von: Lu, Zhufeng, et al.
Veröffentlicht: (2025)
von: Lu, Zhufeng, et al.
Veröffentlicht: (2025)
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
The Blessing and Curse of Dimensionality in Safety Alignment
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
von: Min, Rui, et al.
Veröffentlicht: (2024)
von: Min, Rui, et al.
Veröffentlicht: (2024)
AlphaApollo: A System for Deep Agentic Reasoning
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025) -
Reasoning as an Adaptive Defense for Safety
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025) -
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025) -
Can LLM Safety Be Ensured by Constraining Parameter Regions?
von: Li, Zongmin, et al.
Veröffentlicht: (2026) -
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
von: Zuo, Qian, et al.
Veröffentlicht: (2025)