Combating Adversarial Attacks with Multi-Agent Debate

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chern, Steffi, Fan, Zhen, Liu, Andy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916088591155200
author Chern, Steffi
Fan, Zhen
Liu, Andy
author_facet Chern, Steffi
Fan, Zhen
Liu, Andy
contents While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach proposed to improve the general quality of language model generations is multi-agent debate, where language models self-evaluate through discussion and feedback arXiv:2305.14325. We implement multi-agent debate between current state-of-the-art language models and evaluate models' susceptibility to red team attacks in both single- and multi-agent settings. We find that multi-agent debate can reduce model toxicity when jailbroken or less capable models are forced to debate with non-jailbroken or more capable models. We also find marginal improvements through the general usage of multi-agent interactions. We further perform adversarial prompt content classification via embedding clustering, and analyze the susceptibility of different models to different types of attack topics.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05998
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Combating Adversarial Attacks with Multi-Agent Debate
Chern, Steffi
Fan, Zhen
Liu, Andy
Computation and Language
Artificial Intelligence
While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach proposed to improve the general quality of language model generations is multi-agent debate, where language models self-evaluate through discussion and feedback arXiv:2305.14325. We implement multi-agent debate between current state-of-the-art language models and evaluate models' susceptibility to red team attacks in both single- and multi-agent settings. We find that multi-agent debate can reduce model toxicity when jailbroken or less capable models are forced to debate with non-jailbroken or more capable models. We also find marginal improvements through the general usage of multi-agent interactions. We further perform adversarial prompt content classification via embedding clustering, and analyze the susceptibility of different models to different types of attack topics.
title Combating Adversarial Attacks with Multi-Agent Debate
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.05998