Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An Empirical Study

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Francia, Jerson, Hansen, Derek, Schooley, Ben, Taylor, Matthew, Murray, Shydra, Snow, Greg
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916656335290368
author Francia, Jerson
Hansen, Derek
Schooley, Ben
Taylor, Matthew
Murray, Shydra
Snow, Greg
author_facet Francia, Jerson
Hansen, Derek
Schooley, Ben
Taylor, Matthew
Murray, Shydra
Snow, Greg
contents This paper explores the use of Large Language Models (LLMs) in spear phishing message generation and evaluates their performance compared to human-authored counterparts. Our pilot study examines the effectiveness of smishing (SMS phishing) messages created by GPT-4 and human authors, which have been personalized for willing targets. The targets assessed these messages in a modified ranked-order experiment using a novel methodology we call TRAPD (Threshold Ranking Approach for Personalized Deception). Experiments involved ranking each spear phishing message from most to least convincing, providing qualitative feedback, and guessing which messages were human- or AI-generated. Results show that LLM-generated messages are often perceived as more convincing than those authored by humans, particularly job-related messages. Targets also struggled to distinguish between human- and AI-generated messages. We analyze different criteria the targets used to assess the persuasiveness and source of messages. This study aims to highlight the urgent need for further research and improved countermeasures against personalized AI-enabled social engineering attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An Empirical Study
Francia, Jerson
Hansen, Derek
Schooley, Ben
Taylor, Matthew
Murray, Shydra
Snow, Greg
Computers and Society
Artificial Intelligence
This paper explores the use of Large Language Models (LLMs) in spear phishing message generation and evaluates their performance compared to human-authored counterparts. Our pilot study examines the effectiveness of smishing (SMS phishing) messages created by GPT-4 and human authors, which have been personalized for willing targets. The targets assessed these messages in a modified ranked-order experiment using a novel methodology we call TRAPD (Threshold Ranking Approach for Personalized Deception). Experiments involved ranking each spear phishing message from most to least convincing, providing qualitative feedback, and guessing which messages were human- or AI-generated. Results show that LLM-generated messages are often perceived as more convincing than those authored by humans, particularly job-related messages. Targets also struggled to distinguish between human- and AI-generated messages. We analyze different criteria the targets used to assess the persuasiveness and source of messages. This study aims to highlight the urgent need for further research and improved countermeasures against personalized AI-enabled social engineering attacks.
title Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An Empirical Study
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2406.13049