Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Can, Wu, Rui, Che, Tong, Zhang, Qixin, Peng, Hongwu, Zhao, Jiahui, Wang, Zhenting, Wei, Wenqi, Han, Ligong, Zhang, Zhao, Cao, Yuan, Tang, Ruixiang, Metaxas, Dimitris N.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988549558272
author Jin, Can
Wu, Rui
Che, Tong
Zhang, Qixin
Peng, Hongwu
Zhao, Jiahui
Wang, Zhenting
Wei, Wenqi
Han, Ligong
Zhang, Zhao
Cao, Yuan
Tang, Ruixiang
Metaxas, Dimitris N.
author_facet Jin, Can
Wu, Rui
Che, Tong
Zhang, Qixin
Peng, Hongwu
Zhao, Jiahui
Wang, Zhenting
Wei, Wenqi
Han, Ligong
Zhang, Zhao
Cao, Yuan
Tang, Ruixiang
Metaxas, Dimitris N.
contents Ensuring that Large Language Models (LLMs) adhere to safety principles without refusing benign requests remains a significant challenge. While OpenAI introduces deliberative alignment (DA) to enhance the safety of its o-series models through reasoning over detailed ``code-like'' safety rules, the effectiveness of this approach in open-source LLMs, which typically lack advanced reasoning capabilities, is understudied. In this work, we systematically evaluate the impact of explicitly specifying extensive safety codes versus demonstrating them through illustrative cases. We find that referencing explicit codes inconsistently improves harmlessness and systematically degrades helpfulness, whereas training on case-augmented simple codes yields more robust and generalized safety behaviors. By guiding LLMs with case-augmented reasoning instead of extensive code-like safety rules, we avoid rigid adherence to narrowly enumerated rules and enable broader adaptability. Building on these insights, we propose CADA, a case-augmented deliberative alignment method for LLMs utilizing reinforcement learning on self-generated safety reasoning chains. CADA effectively enhances harmlessness, improves robustness against attacks, and reduces over-refusal while preserving utility across diverse benchmarks, offering a practical alternative to rule-only DA for improving safety while maintaining helpfulness.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08000
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
Jin, Can
Wu, Rui
Che, Tong
Zhang, Qixin
Peng, Hongwu
Zhao, Jiahui
Wang, Zhenting
Wei, Wenqi
Han, Ligong
Zhang, Zhao
Cao, Yuan
Tang, Ruixiang
Metaxas, Dimitris N.
Artificial Intelligence
Software Engineering
Ensuring that Large Language Models (LLMs) adhere to safety principles without refusing benign requests remains a significant challenge. While OpenAI introduces deliberative alignment (DA) to enhance the safety of its o-series models through reasoning over detailed ``code-like'' safety rules, the effectiveness of this approach in open-source LLMs, which typically lack advanced reasoning capabilities, is understudied. In this work, we systematically evaluate the impact of explicitly specifying extensive safety codes versus demonstrating them through illustrative cases. We find that referencing explicit codes inconsistently improves harmlessness and systematically degrades helpfulness, whereas training on case-augmented simple codes yields more robust and generalized safety behaviors. By guiding LLMs with case-augmented reasoning instead of extensive code-like safety rules, we avoid rigid adherence to narrowly enumerated rules and enable broader adaptability. Building on these insights, we propose CADA, a case-augmented deliberative alignment method for LLMs utilizing reinforcement learning on self-generated safety reasoning chains. CADA effectively enhances harmlessness, improves robustness against attacks, and reduces over-refusal while preserving utility across diverse benchmarks, offering a practical alternative to rule-only DA for improving safety while maintaining helpfulness.
title Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2601.08000