SafeLawBench: Towards Safe Alignment of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Chuxue, Zhu, Han, Ji, Jiaming, Sun, Qichao, Zhu, Zhenghao, Wu, Yinyu, Dai, Juntao, Yang, Yaodong, Han, Sirui, Guo, Yike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909641642868736
author Cao, Chuxue
Zhu, Han
Ji, Jiaming
Sun, Qichao
Zhu, Zhenghao
Wu, Yinyu
Dai, Juntao
Yang, Yaodong
Han, Sirui
Guo, Yike
author_facet Cao, Chuxue
Zhu, Han
Ji, Jiaming
Sun, Qichao
Zhu, Zhenghao
Wu, Yinyu
Dai, Juntao
Yang, Yaodong
Han, Sirui
Guo, Yike
contents With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns. However, there is still a lack of definitive standards for evaluating their safety due to the subjective nature of current safety benchmarks. To address this gap, we conducted the first exploration of LLMs' safety evaluation from a legal perspective by proposing the SafeLawBench benchmark. SafeLawBench categorizes safety risks into three levels based on legal standards, providing a systematic and comprehensive framework for evaluation. It comprises 24,860 multi-choice questions and 1,106 open-domain question-answering (QA) tasks. Our evaluation included 2 closed-source LLMs and 18 open-source LLMs using zero-shot and few-shot prompting, highlighting the safety features of each model. We also evaluated the LLMs' safety-related reasoning stability and refusal behavior. Additionally, we found that a majority voting mechanism can enhance model performance. Notably, even leading SOTA models like Claude-3.5-Sonnet and GPT-4o have not exceeded 80.5% accuracy in multi-choice tasks on SafeLawBench, while the average accuracy of 20 LLMs remains at 68.8\%. We urge the community to prioritize research on the safety of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06636
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeLawBench: Towards Safe Alignment of Large Language Models
Cao, Chuxue
Zhu, Han
Ji, Jiaming
Sun, Qichao
Zhu, Zhenghao
Wu, Yinyu
Dai, Juntao
Yang, Yaodong
Han, Sirui
Guo, Yike
Computation and Language
With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns. However, there is still a lack of definitive standards for evaluating their safety due to the subjective nature of current safety benchmarks. To address this gap, we conducted the first exploration of LLMs' safety evaluation from a legal perspective by proposing the SafeLawBench benchmark. SafeLawBench categorizes safety risks into three levels based on legal standards, providing a systematic and comprehensive framework for evaluation. It comprises 24,860 multi-choice questions and 1,106 open-domain question-answering (QA) tasks. Our evaluation included 2 closed-source LLMs and 18 open-source LLMs using zero-shot and few-shot prompting, highlighting the safety features of each model. We also evaluated the LLMs' safety-related reasoning stability and refusal behavior. Additionally, we found that a majority voting mechanism can enhance model performance. Notably, even leading SOTA models like Claude-3.5-Sonnet and GPT-4o have not exceeded 80.5% accuracy in multi-choice tasks on SafeLawBench, while the average accuracy of 20 LLMs remains at 68.8\%. We urge the community to prioritize research on the safety of LLMs.
title SafeLawBench: Towards Safe Alignment of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.06636