Dr. Jekyll and Mr. Hyde: Two Faces of LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Collu, Matteo Gioele, Janssen-Groesbeek, Tom, Koffas, Stefanos, Conti, Mauro, Picek, Stjepan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911633616404480
author Collu, Matteo Gioele
Janssen-Groesbeek, Tom
Koffas, Stefanos
Conti, Mauro
Picek, Stjepan
author_facet Collu, Matteo Gioele
Janssen-Groesbeek, Tom
Koffas, Stefanos
Conti, Mauro
Picek, Stjepan
contents Large Language Models (LLMs) are being integrated into applications such as chatbots or email assistants. To prevent improper responses, safety mechanisms, such as Reinforcement Learning from Human Feedback (RLHF), are implemented in them. In this work, we bypass these safety measures for ChatGPT, Gemini, and Deepseek by making them impersonate complex personas with personality characteristics that are not aligned with a truthful assistant. First, we create elaborate biographies of these personas, which we then use in a new session with the same chatbots. Our conversations then follow a role-play style to elicit prohibited responses. Using personas, we show that prohibited responses are provided, making it possible to obtain unauthorized, illegal, or harmful information when querying ChatGPT, Gemini, and Deepseek. We show that these chatbots are vulnerable to this attack by getting dangerous information for 40 out of 40 illicit questions in GPT-4.1-mini, Gemini-1.5-flash, 39 out of 40 in GPT-4o-mini, 38 out of 40 in GPT-3.5-turbo, and 2 out of 2 cases in Gemini-2.5-flash and DeepSeek V3. The attack can be carried out manually or automatically using a support LLM, and has proven effective against models deployed between 2023 and 2025.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03853
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
Collu, Matteo Gioele
Janssen-Groesbeek, Tom
Koffas, Stefanos
Conti, Mauro
Picek, Stjepan
Cryptography and Security
Machine Learning
K.6.5; D.4.6; I.2.6; K.6.5
Large Language Models (LLMs) are being integrated into applications such as chatbots or email assistants. To prevent improper responses, safety mechanisms, such as Reinforcement Learning from Human Feedback (RLHF), are implemented in them. In this work, we bypass these safety measures for ChatGPT, Gemini, and Deepseek by making them impersonate complex personas with personality characteristics that are not aligned with a truthful assistant. First, we create elaborate biographies of these personas, which we then use in a new session with the same chatbots. Our conversations then follow a role-play style to elicit prohibited responses. Using personas, we show that prohibited responses are provided, making it possible to obtain unauthorized, illegal, or harmful information when querying ChatGPT, Gemini, and Deepseek. We show that these chatbots are vulnerable to this attack by getting dangerous information for 40 out of 40 illicit questions in GPT-4.1-mini, Gemini-1.5-flash, 39 out of 40 in GPT-4o-mini, 38 out of 40 in GPT-3.5-turbo, and 2 out of 2 cases in Gemini-2.5-flash and DeepSeek V3. The attack can be carried out manually or automatically using a support LLM, and has proven effective against models deployed between 2023 and 2025.
title Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
topic Cryptography and Security
Machine Learning
K.6.5; D.4.6; I.2.6; K.6.5
url https://arxiv.org/abs/2312.03853