DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Andrew, Xu, Quentin, Lin, Matthieu, Wang, Shenzhi, Liu, Yong-jin, Zheng, Zilong, Huang, Gao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909435517992960
author Zhao, Andrew
Xu, Quentin
Lin, Matthieu
Wang, Shenzhi
Liu, Yong-jin
Zheng, Zilong
Huang, Gao
author_facet Zhao, Andrew
Xu, Quentin
Lin, Matthieu
Wang, Shenzhi
Liu, Yong-jin
Zheng, Zilong
Huang, Gao
contents Recent advances in large language model assistants have made them indispensable, raising significant concerns over managing their safety. Automated red teaming offers a promising alternative to the labor-intensive and error-prone manual probing for vulnerabilities, providing more consistent and scalable safety evaluations. However, existing approaches often compromise diversity by focusing on maximizing attack success rate. Additionally, methods that decrease the cosine similarity from historical embeddings with semantic diversity rewards lead to novelty stagnation as history grows. To address these issues, we introduce DiveR-CT, which relaxes conventional constraints on the objective and semantic reward, granting greater freedom for the policy to enhance diversity. Our experiments demonstrate DiveR-CT's marked superiority over baselines by 1) generating data that perform better in various diversity metrics across different attack success rate levels, 2) better-enhancing resiliency in blue team models through safety tuning based on collected data, 3) allowing dynamic control of objective weights for reliable and controllable attack success rates, and 4) reducing susceptibility to reward overoptimization. Overall, our method provides an effective and efficient approach to LLM red teaming, accelerating real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19026
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
Zhao, Andrew
Xu, Quentin
Lin, Matthieu
Wang, Shenzhi
Liu, Yong-jin
Zheng, Zilong
Huang, Gao
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Recent advances in large language model assistants have made them indispensable, raising significant concerns over managing their safety. Automated red teaming offers a promising alternative to the labor-intensive and error-prone manual probing for vulnerabilities, providing more consistent and scalable safety evaluations. However, existing approaches often compromise diversity by focusing on maximizing attack success rate. Additionally, methods that decrease the cosine similarity from historical embeddings with semantic diversity rewards lead to novelty stagnation as history grows. To address these issues, we introduce DiveR-CT, which relaxes conventional constraints on the objective and semantic reward, granting greater freedom for the policy to enhance diversity. Our experiments demonstrate DiveR-CT's marked superiority over baselines by 1) generating data that perform better in various diversity metrics across different attack success rate levels, 2) better-enhancing resiliency in blue team models through safety tuning based on collected data, 3) allowing dynamic control of objective weights for reliable and controllable attack success rates, and 4) reducing susceptibility to reward overoptimization. Overall, our method provides an effective and efficient approach to LLM red teaming, accelerating real-world deployment.
title DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2405.19026