SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Tianhao, Lu, Jingyu, Chu, Chuangxin, Zeng, Tianyu, Zheng, Yujia, Li, Mei, Huang, Haotian, Wu, Bin, Liu, Zuoxian, Ma, Kai, Yuan, Xuejing, Wang, Xingkai, Ding, Keyan, Chen, Huajun, Zhang, Qiang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929631980945408
author Li, Tianhao
Lu, Jingyu
Chu, Chuangxin
Zeng, Tianyu
Zheng, Yujia
Li, Mei
Huang, Haotian
Wu, Bin
Liu, Zuoxian
Ma, Kai
Yuan, Xuejing
Wang, Xingkai
Ding, Keyan
Chen, Huajun
Zhang, Qiang
author_facet Li, Tianhao
Lu, Jingyu
Chu, Chuangxin
Zeng, Tianyu
Zheng, Yujia
Li, Mei
Huang, Haotian
Wu, Bin
Liu, Zuoxian
Ma, Kai
Yuan, Xuejing
Wang, Xingkai
Ding, Keyan
Chen, Huajun
Zhang, Qiang
contents Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focusing on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchmark designed to evaluate the safety alignment of LLMs across a range of scientific tasks. SciSafeEval spans multiple scientific languages-including textual, molecular, protein, and genomic-and covers a wide range of scientific domains. We evaluate LLMs in zero-shot, few-shot and chain-of-thought settings, and introduce a "jailbreak" enhancement feature that challenges LLMs equipped with safety guardrails, rigorously testing their defenses against malicious intention. Our benchmark surpasses existing safety datasets in both scale and scope, providing a robust platform for assessing the safety and performance of LLMs in scientific contexts. This work aims to facilitate the responsible development and deployment of LLMs, promoting alignment with safety and ethical standards in scientific research.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03769
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
Li, Tianhao
Lu, Jingyu
Chu, Chuangxin
Zeng, Tianyu
Zheng, Yujia
Li, Mei
Huang, Haotian
Wu, Bin
Liu, Zuoxian
Ma, Kai
Yuan, Xuejing
Wang, Xingkai
Ding, Keyan
Chen, Huajun
Zhang, Qiang
Computation and Language
Artificial Intelligence
Cryptography and Security
Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focusing on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchmark designed to evaluate the safety alignment of LLMs across a range of scientific tasks. SciSafeEval spans multiple scientific languages-including textual, molecular, protein, and genomic-and covers a wide range of scientific domains. We evaluate LLMs in zero-shot, few-shot and chain-of-thought settings, and introduce a "jailbreak" enhancement feature that challenges LLMs equipped with safety guardrails, rigorously testing their defenses against malicious intention. Our benchmark surpasses existing safety datasets in both scale and scope, providing a robust platform for assessing the safety and performance of LLMs in scientific contexts. This work aims to facilitate the responsible development and deployment of LLMs, promoting alignment with safety and ethical standards in scientific research.
title SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2410.03769