Robustness of Large Language Models Against Adversarial Attacks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tao, Yiyi, Shen, Yixian, Zhang, Hang, Shen, Yanxin, Wang, Lun, Shi, Chuanqi, Du, Shaoshuai
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912166172426240
author Tao, Yiyi
Shen, Yixian
Zhang, Hang
Shen, Yanxin
Wang, Lun
Shi, Chuanqi
Du, Shaoshuai
author_facet Tao, Yiyi
Shen, Yixian
Zhang, Hang
Shen, Yanxin
Wang, Lun
Shi, Chuanqi
Du, Shaoshuai
contents The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17011
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robustness of Large Language Models Against Adversarial Attacks
Tao, Yiyi
Shen, Yixian
Zhang, Hang
Shen, Yanxin
Wang, Lun
Shi, Chuanqi
Du, Shaoshuai
Computation and Language
The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs.
title Robustness of Large Language Models Against Adversarial Attacks
topic Computation and Language
url https://arxiv.org/abs/2412.17011