SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kumar, Anurakt, Kumar, Divyanshu, Loya, Jatan, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910573184155648
author Kumar, Anurakt
Kumar, Divyanshu
Loya, Jatan
Birur, Nitin Aravind
Baswa, Tanay
Agarwal, Sahil
Harshangi, Prashanth
author_facet Kumar, Anurakt
Kumar, Divyanshu
Loya, Jatan
Birur, Nitin Aravind
Baswa, Tanay
Agarwal, Sahil
Harshangi, Prashanth
contents We introduce Synthetic Alignment data Generation for Safety Evaluation and Red Teaming (SAGE-RT or SAGE) a novel pipeline for generating synthetic alignment and red-teaming data. Existing methods fall short in creating nuanced and diverse datasets, providing necessary control over the data generation and validation processes, or require large amount of manually generated seed data. SAGE addresses these limitations by using a detailed taxonomy to produce safety-alignment and red-teaming data across a wide range of topics. We generated 51,000 diverse and in-depth prompt-response pairs, encompassing over 1,500 topics of harmfulness and covering variations of the most frequent types of jailbreaking prompts faced by large language models (LLMs). We show that the red-teaming data generated through SAGE jailbreaks state-of-the-art LLMs in more than 27 out of 32 sub-categories, and in more than 58 out of 279 leaf-categories (sub-sub categories). The attack success rate for GPT-4o, GPT-3.5-turbo is 100% over the sub-categories of harmfulness. Our approach avoids the pitfalls of synthetic safety-training data generation such as mode collapse and lack of nuance in the generation pipeline by ensuring a detailed coverage of harmful topics using iterative expansion of the topics and conditioning the outputs on the generated raw-text. This method can be used to generate red-teaming and alignment data for LLM Safety completely synthetically to make LLMs safer or for red-teaming the models over a diverse range of topics.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
Kumar, Anurakt
Kumar, Divyanshu
Loya, Jatan
Birur, Nitin Aravind
Baswa, Tanay
Agarwal, Sahil
Harshangi, Prashanth
Artificial Intelligence
Computation and Language
Cryptography and Security
We introduce Synthetic Alignment data Generation for Safety Evaluation and Red Teaming (SAGE-RT or SAGE) a novel pipeline for generating synthetic alignment and red-teaming data. Existing methods fall short in creating nuanced and diverse datasets, providing necessary control over the data generation and validation processes, or require large amount of manually generated seed data. SAGE addresses these limitations by using a detailed taxonomy to produce safety-alignment and red-teaming data across a wide range of topics. We generated 51,000 diverse and in-depth prompt-response pairs, encompassing over 1,500 topics of harmfulness and covering variations of the most frequent types of jailbreaking prompts faced by large language models (LLMs). We show that the red-teaming data generated through SAGE jailbreaks state-of-the-art LLMs in more than 27 out of 32 sub-categories, and in more than 58 out of 279 leaf-categories (sub-sub categories). The attack success rate for GPT-4o, GPT-3.5-turbo is 100% over the sub-categories of harmfulness. Our approach avoids the pitfalls of synthetic safety-training data generation such as mode collapse and lack of nuance in the generation pipeline by ensuring a detailed coverage of harmful topics using iterative expansion of the topics and conditioning the outputs on the generated raw-text. This method can be used to generate red-teaming and alignment data for LLM Safety completely synthetically to make LLMs safer or for red-teaming the models over a diverse range of topics.
title SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
topic Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2408.11851