Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Xikang, Tang, Xuehai, Hu, Songlin, Han, Jizhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916240076832768
author Yang, Xikang
Tang, Xuehai
Hu, Songlin
Han, Jizhong
author_facet Yang, Xikang
Tang, Xuehai
Hu, Songlin
Han, Jizhong
contents Large language models (LLMs) have achieved remarkable performance in various natural language processing tasks, especially in dialogue systems. However, LLM may also pose security and moral threats, especially in multi round conversations where large models are more easily guided by contextual content, resulting in harmful or biased responses. In this paper, we present a novel method to attack LLMs in multi-turn dialogues, called CoA (Chain of Attack). CoA is a semantic-driven contextual multi-turn attack method that adaptively adjusts the attack policy through contextual feedback and semantic relevance during multi-turn of dialogue with a large model, resulting in the model producing unreasonable or harmful content. We evaluate CoA on different LLMs and datasets, and show that it can effectively expose the vulnerabilities of LLMs, and outperform existing attack methods. Our work provides a new perspective and tool for attacking and defending LLMs, and contributes to the security and ethical assessment of dialogue systems.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05610
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
Yang, Xikang
Tang, Xuehai
Hu, Songlin
Han, Jizhong
Computation and Language
Cryptography and Security
Machine Learning
Large language models (LLMs) have achieved remarkable performance in various natural language processing tasks, especially in dialogue systems. However, LLM may also pose security and moral threats, especially in multi round conversations where large models are more easily guided by contextual content, resulting in harmful or biased responses. In this paper, we present a novel method to attack LLMs in multi-turn dialogues, called CoA (Chain of Attack). CoA is a semantic-driven contextual multi-turn attack method that adaptively adjusts the attack policy through contextual feedback and semantic relevance during multi-turn of dialogue with a large model, resulting in the model producing unreasonable or harmful content. We evaluate CoA on different LLMs and datasets, and show that it can effectively expose the vulnerabilities of LLMs, and outperform existing attack methods. Our work provides a new perspective and tool for attacking and defending LLMs, and contributes to the security and ethical assessment of dialogue systems.
title Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2405.05610