Lifelong Safety Alignment for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haoyu, Qin, Zeyu, Zhao, Yifei, Du, Chao, Lin, Min, Wang, Xueqian, Pang, Tianyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910969129598976
author Wang, Haoyu
Qin, Zeyu
Zhao, Yifei
Du, Chao
Lin, Min
Wang, Xueqian
Pang, Tianyu
author_facet Wang, Haoyu
Qin, Zeyu
Zhao, Yifei
Du, Chao
Lin, Min
Wang, Xueqian
Pang, Tianyu
contents LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for unseen attacks that may arise during deployment. To address this, we propose a lifelong safety alignment framework that enables LLMs to continuously adapt to new and evolving jailbreaking strategies. Our framework introduces a competitive setup between two components: a Meta-Attacker, trained to actively discover novel jailbreaking strategies, and a Defender, trained to resist them. To effectively warm up the Meta-Attacker, we first leverage the GPT-4o API to extract key insights from a large collection of jailbreak-related research papers. Through iterative training, the first iteration Meta-Attacker achieves a 73% attack success rate (ASR) on RR and a 57% transfer ASR on LAT using only single-turn attacks. Meanwhile, the Defender progressively improves its robustness and ultimately reduces the Meta-Attacker's success rate to just 7%, enabling safer and more reliable deployment of LLMs in open-ended environments. The code is available at https://github.com/sail-sg/LifelongSafetyAlignment.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20259
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lifelong Safety Alignment for Language Models
Wang, Haoyu
Qin, Zeyu
Zhao, Yifei
Du, Chao
Lin, Min
Wang, Xueqian
Pang, Tianyu
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for unseen attacks that may arise during deployment. To address this, we propose a lifelong safety alignment framework that enables LLMs to continuously adapt to new and evolving jailbreaking strategies. Our framework introduces a competitive setup between two components: a Meta-Attacker, trained to actively discover novel jailbreaking strategies, and a Defender, trained to resist them. To effectively warm up the Meta-Attacker, we first leverage the GPT-4o API to extract key insights from a large collection of jailbreak-related research papers. Through iterative training, the first iteration Meta-Attacker achieves a 73% attack success rate (ASR) on RR and a 57% transfer ASR on LAT using only single-turn attacks. Meanwhile, the Defender progressively improves its robustness and ultimately reduces the Meta-Attacker's success rate to just 7%, enabling safer and more reliable deployment of LLMs in open-ended environments. The code is available at https://github.com/sail-sg/LifelongSafetyAlignment.
title Lifelong Safety Alignment for Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.20259