Robustness of Large Language Models Against Adversarial Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912166172426240 |
|---|---|
| author | Tao, Yiyi Shen, Yixian Zhang, Hang Shen, Yanxin Wang, Lun Shi, Chuanqi Du, Shaoshuai |
| author_facet | Tao, Yiyi Shen, Yixian Zhang, Hang Shen, Yanxin Wang, Lun Shi, Chuanqi Du, Shaoshuai |
| contents | The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_17011 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Robustness of Large Language Models Against Adversarial Attacks Tao, Yiyi Shen, Yixian Zhang, Hang Shen, Yanxin Wang, Lun Shi, Chuanqi Du, Shaoshuai Computation and Language The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs. |
| title | Robustness of Large Language Models Against Adversarial Attacks |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.17011 |