Promoting Online Safety by Simulating Unsafe Conversations with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoffman, Owen, Peng, Kangze, You, Zehua, Kamal, Sajid, Venkatagiri, Sukrit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909712106127360
author Hoffman, Owen
Peng, Kangze
You, Zehua
Kamal, Sajid
Venkatagiri, Sukrit
author_facet Hoffman, Owen
Peng, Kangze
You, Zehua
Kamal, Sajid
Venkatagiri, Sukrit
contents Generative AI, including large language models (LLMs) have the potential -- and already are being used -- to increase the speed, scale, and types of unsafe conversations online. LLMs lower the barrier for entry for bad actors to create unsafe conversations in particular because of their ability to generate persuasive and human-like text. In our current work, we explore ways to promote online safety by teaching people about unsafe conversations that can occur online with and without LLMs. We build on prior work that shows that LLMs can successfully simulate scam conversations. We also leverage research in the learning sciences that shows that providing feedback on one's hypothetical actions can promote learning. In particular, we focus on simulating scam conversations using LLMs. Our work incorporates two LLMs that converse with each other to simulate realistic, unsafe conversations that people may encounter online between a scammer LLM and a target LLM but users of our system are asked provide feedback to the target LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Promoting Online Safety by Simulating Unsafe Conversations with LLMs
Hoffman, Owen
Peng, Kangze
You, Zehua
Kamal, Sajid
Venkatagiri, Sukrit
Human-Computer Interaction
Artificial Intelligence
Generative AI, including large language models (LLMs) have the potential -- and already are being used -- to increase the speed, scale, and types of unsafe conversations online. LLMs lower the barrier for entry for bad actors to create unsafe conversations in particular because of their ability to generate persuasive and human-like text. In our current work, we explore ways to promote online safety by teaching people about unsafe conversations that can occur online with and without LLMs. We build on prior work that shows that LLMs can successfully simulate scam conversations. We also leverage research in the learning sciences that shows that providing feedback on one's hypothetical actions can promote learning. In particular, we focus on simulating scam conversations using LLMs. Our work incorporates two LLMs that converse with each other to simulate realistic, unsafe conversations that people may encounter online between a scammer LLM and a target LLM but users of our system are asked provide feedback to the target LLM.
title Promoting Online Safety by Simulating Unsafe Conversations with LLMs
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2507.22267