Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908625552801792 |
|---|---|
| author | Chang, Chen-Wei Sarkar, Shailik Mitra, Shutonu Zhang, Qi Salemi, Hossein Purohit, Hemant Zhang, Fengxiu Hong, Michin Cho, Jin-Hee Lu, Chang-Tien |
| author_facet | Chang, Chen-Wei Sarkar, Shailik Mitra, Shutonu Zhang, Qi Salemi, Hossein Purohit, Hemant Zhang, Fengxiu Hong, Michin Cho, Jin-Hee Lu, Chang-Tien |
| contents | Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_00621 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance Chang, Chen-Wei Sarkar, Shailik Mitra, Shutonu Zhang, Qi Salemi, Hossein Purohit, Hemant Zhang, Fengxiu Hong, Michin Cho, Jin-Hee Lu, Chang-Tien Cryptography and Security Artificial Intelligence Computers and Society Social and Information Networks Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness. |
| title | Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance |
| topic | Cryptography and Security Artificial Intelligence Computers and Society Social and Information Networks |
| url | https://arxiv.org/abs/2412.00621 |