Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Can, Wu, Rui, Che, Tong, Zhang, Qixin, Peng, Hongwu, Zhao, Jiahui, Wang, Zhenting, Wei, Wenqi, Han, Ligong, Zhang, Zhao, Cao, Yuan, Tang, Ruixiang, Metaxas, Dimitris N. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
by: Xu, Qipan, et al.
Published: (2025)
by: Xu, Qipan, et al.
Published: (2025)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
Data Augmentation for High-Fidelity Generation of CAR-T/NK Immunological Synapse Images
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
RankFlow: A Multi-Role Collaborative Reranking Workflow Utilizing Large Language Models
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
DORY: Deliberative Prompt Recovery for LLM
by: Gao, Lirong, et al.
Published: (2024)
by: Gao, Lirong, et al.
Published: (2024)
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
by: Jiang, Jinhao, et al.
Published: (2024)
by: Jiang, Jinhao, et al.
Published: (2024)
Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
by: Wu, Rui, et al.
Published: (2025)
by: Wu, Rui, et al.
Published: (2025)
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
by: Wang, Zhenting, et al.
Published: (2023)
by: Wang, Zhenting, et al.
Published: (2023)
PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation
by: Mo, Wenyi, et al.
Published: (2025)
by: Mo, Wenyi, et al.
Published: (2025)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Evidence Over Plans: Online Trajectory Verification for Skill Distillation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
Deliberative Dynamics and Value Alignment in LLM Debates
by: Sachdeva, Pratik S., et al.
Published: (2025)
by: Sachdeva, Pratik S., et al.
Published: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
Meaningless Tokens, Meaningful Gains: How Activation Shifts Enhance LLM Reasoning
by: Shi, Zeru, et al.
Published: (2025)
by: Shi, Zeru, et al.
Published: (2025)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
by: Li, Jiakang, et al.
Published: (2026)
by: Li, Jiakang, et al.
Published: (2026)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
by: Zhao, Libang, et al.
Published: (2026)
by: Zhao, Libang, et al.
Published: (2026)
Token-Budget-Aware LLM Reasoning
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
Reason4Rec: Large Language Models for Recommendation with Deliberative User Preference Alignment
by: Fang, Yi, et al.
Published: (2025)
by: Fang, Yi, et al.
Published: (2025)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
by: Zhang, Tunyu, et al.
Published: (2025)
by: Zhang, Tunyu, et al.
Published: (2025)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
by: Vega, Jason, et al.
Published: (2024)
by: Vega, Jason, et al.
Published: (2024)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
by: Mou, Yutao, et al.
Published: (2025)
by: Mou, Yutao, et al.
Published: (2025)
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
by: Zhao, Shiyu, et al.
Published: (2024)
by: Zhao, Shiyu, et al.
Published: (2024)
Similar Items
-
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025) -
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
by: Jin, Can, et al.
Published: (2024) -
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
by: Jin, Can, et al.
Published: (2025) -
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025) -
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
by: Jin, Can, et al.
Published: (2024)