GuardReasoner: Towards Reasoning-based LLM Safeguards

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yue, Gao, Hongcheng, Zhai, Shengfang, He, Yufei, Xia, Jun, Hu, Zhengyu, Chen, Yulin, Yang, Xihong, Zhang, Jiaheng, Li, Stan Z., Xiong, Hui, Hooi, Bryan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915558964854784
author Liu, Yue
Gao, Hongcheng
Zhai, Shengfang
He, Yufei
Xia, Jun
Hu, Zhengyu
Chen, Yulin
Yang, Xihong
Zhang, Jiaheng
Li, Stan Z.
Xiong, Hui
Hooi, Bryan
author_facet Liu, Yue
Gao, Hongcheng
Zhai, Shengfang
He, Yufei
Xia, Jun
Hu, Zhengyu
Chen, Yulin
Yang, Xihong
Zhang, Jiaheng
Li, Stan Z.
Xiong, Hui
Hooi, Bryan
contents As LLMs increasingly impact safety-critical applications, ensuring their safety using guardrails remains a key challenge. This paper proposes GuardReasoner, a new safeguard for LLMs, by guiding the guard model to learn to reason. Concretely, we first create the GuardReasonerTrain dataset, which consists of 127K samples with 460K detailed reasoning steps. Then, we introduce reasoning SFT to unlock the reasoning capability of guard models. In addition, we present hard sample DPO to further strengthen their reasoning ability. In this manner, GuardReasoner achieves better performance, explainability, and generalizability. Extensive experiments and analyses on 13 benchmarks of 3 guardrail tasks demonstrate its superiority. Remarkably, GuardReasoner 8B surpasses GPT-4o+CoT by 5.74% and LLaMA Guard 3 8B by 20.84% F1 score on average. We release the training data, code, and models with different scales (1B, 3B, 8B) of GuardReasoner : https://github.com/yueliu1999/GuardReasoner/.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GuardReasoner: Towards Reasoning-based LLM Safeguards
Liu, Yue
Gao, Hongcheng
Zhai, Shengfang
He, Yufei
Xia, Jun
Hu, Zhengyu
Chen, Yulin
Yang, Xihong
Zhang, Jiaheng
Li, Stan Z.
Xiong, Hui
Hooi, Bryan
Cryptography and Security
Artificial Intelligence
Machine Learning
As LLMs increasingly impact safety-critical applications, ensuring their safety using guardrails remains a key challenge. This paper proposes GuardReasoner, a new safeguard for LLMs, by guiding the guard model to learn to reason. Concretely, we first create the GuardReasonerTrain dataset, which consists of 127K samples with 460K detailed reasoning steps. Then, we introduce reasoning SFT to unlock the reasoning capability of guard models. In addition, we present hard sample DPO to further strengthen their reasoning ability. In this manner, GuardReasoner achieves better performance, explainability, and generalizability. Extensive experiments and analyses on 13 benchmarks of 3 guardrail tasks demonstrate its superiority. Remarkably, GuardReasoner 8B surpasses GPT-4o+CoT by 5.74% and LLaMA Guard 3 8B by 20.84% F1 score on average. We release the training data, code, and models with different scales (1B, 3B, 8B) of GuardReasoner : https://github.com/yueliu1999/GuardReasoner/.
title GuardReasoner: Towards Reasoning-based LLM Safeguards
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.18492