Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Chen-Wei, Sarkar, Shailik, Mitra, Shutonu, Zhang, Qi, Salemi, Hossein, Purohit, Hemant, Zhang, Fengxiu, Hong, Michin, Cho, Jin-Hee, Lu, Chang-Tien
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908625552801792
author Chang, Chen-Wei
Sarkar, Shailik
Mitra, Shutonu
Zhang, Qi
Salemi, Hossein
Purohit, Hemant
Zhang, Fengxiu
Hong, Michin
Cho, Jin-Hee
Lu, Chang-Tien
author_facet Chang, Chen-Wei
Sarkar, Shailik
Mitra, Shutonu
Zhang, Qi
Salemi, Hossein
Purohit, Hemant
Zhang, Fengxiu
Hong, Michin
Cho, Jin-Hee
Lu, Chang-Tien
contents Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
Chang, Chen-Wei
Sarkar, Shailik
Mitra, Shutonu
Zhang, Qi
Salemi, Hossein
Purohit, Hemant
Zhang, Fengxiu
Hong, Michin
Cho, Jin-Hee
Lu, Chang-Tien
Cryptography and Security
Artificial Intelligence
Computers and Society
Social and Information Networks
Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness.
title Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
topic Cryptography and Security
Artificial Intelligence
Computers and Society
Social and Information Networks
url https://arxiv.org/abs/2412.00621