The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toth, Rebeka, Bisztray, Tamas, Gruschka, Nils
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912975693021184
author Toth, Rebeka
Bisztray, Tamas
Gruschka, Nils
author_facet Toth, Rebeka
Bisztray, Tamas
Gruschka, Nils
contents In this paper, we introduce a metadata-enriched generation framework (PhishFuzzer) that seeds real emails into Large Language Models (LLMs) to produce 23,100 diverse, structurally consistent email variants across controlled entity and length dimensions. Unlike prior corpora, our dataset features strict three-class labels (Phishing, Spam, Valid), provides full URL and attachment metadata, and annotates each email with attacker intent. Using this dataset, we benchmark two state-of-the-art LLMs (Qwen-2.5-72B and Gemini-3.1-Pro) under both Basic (body, subject) and Full (+URL, sender, attachment) settings. By applying formal confidence metrics (Task Success Rate and Confidence Index), we analyze model reliability, robustness against linguistic fuzzing, and the impact of structural metadata on detection accuracy. Our fully open-source framework and dataset provide a rigorous foundation for evaluating next-generation email security systems. To support open science, we make the PhishFuzzer Dataset, the generation scripts and prompts available on GitHub: https://github.com/DataPhish/PhishFuzzer
format Preprint
id arxiv_https___arxiv_org_abs_2511_21448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
Toth, Rebeka
Bisztray, Tamas
Gruschka, Nils
Cryptography and Security
Artificial Intelligence
Databases
In this paper, we introduce a metadata-enriched generation framework (PhishFuzzer) that seeds real emails into Large Language Models (LLMs) to produce 23,100 diverse, structurally consistent email variants across controlled entity and length dimensions. Unlike prior corpora, our dataset features strict three-class labels (Phishing, Spam, Valid), provides full URL and attachment metadata, and annotates each email with attacker intent. Using this dataset, we benchmark two state-of-the-art LLMs (Qwen-2.5-72B and Gemini-3.1-Pro) under both Basic (body, subject) and Full (+URL, sender, attachment) settings. By applying formal confidence metrics (Task Success Rate and Confidence Index), we analyze model reliability, robustness against linguistic fuzzing, and the impact of structural metadata on detection accuracy. Our fully open-source framework and dataset provide a rigorous foundation for evaluating next-generation email security systems. To support open science, we make the PhishFuzzer Dataset, the generation scripts and prompts available on GitHub: https://github.com/DataPhish/PhishFuzzer
title The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
topic Cryptography and Security
Artificial Intelligence
Databases
url https://arxiv.org/abs/2511.21448