A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Zhai, Keke
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913632178143232
author Zhai, Keke
author_facet Zhai, Keke
contents Currently, large models are prone to generating harmful content when faced with complex attack instructions, significantly reducing their defensive capabilities. To address this issue, this paper proposes a method based on constructing data aligned with multi-dimensional attack defense to enhance the generative security of large models. The core of our method lies in improving the effectiveness of safe alignment learning for large models by innova-tively increasing the diversity of attack instruction dimensions and the accuracy of generat-ing safe responses. To validate the effectiveness of our method, beyond existing security evaluation benchmarks, we additionally designed new security evaluation benchmarks and conducted comparative experiments using Llama3.2 as the baseline model. The final ex-perimental results demonstrate that our method can significantly improve the generative security of large models under complex instructional attacks, while also maintaining and enhancing the models' general capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00517
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
Zhai, Keke
Cryptography and Security
Artificial Intelligence
Currently, large models are prone to generating harmful content when faced with complex attack instructions, significantly reducing their defensive capabilities. To address this issue, this paper proposes a method based on constructing data aligned with multi-dimensional attack defense to enhance the generative security of large models. The core of our method lies in improving the effectiveness of safe alignment learning for large models by innova-tively increasing the diversity of attack instruction dimensions and the accuracy of generat-ing safe responses. To validate the effectiveness of our method, beyond existing security evaluation benchmarks, we additionally designed new security evaluation benchmarks and conducted comparative experiments using Llama3.2 as the baseline model. The final ex-perimental results demonstrate that our method can significantly improve the generative security of large models under complex instructional attacks, while also maintaining and enhancing the models' general capabilities.
title A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2501.00517