PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dalton, Thomas, Gowda, Hemanth, Rao, Girish, Pargi, Sachin, Khodabakhshi, Alireza Hadj, Rombs, Joseph, Jou, Stephan, Marwah, Manish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917268561068032
author Dalton, Thomas
Gowda, Hemanth
Rao, Girish
Pargi, Sachin
Khodabakhshi, Alireza Hadj
Rombs, Joseph
Jou, Stephan
Marwah, Manish
author_facet Dalton, Thomas
Gowda, Hemanth
Rao, Girish
Pargi, Sachin
Khodabakhshi, Alireza Hadj
Rombs, Joseph
Jou, Stephan
Marwah, Manish
contents Phishing remains a pervasive and growing threat, inflicting heavy economic and reputational damage. While machine learning has been effective in real-time detection of phishing attacks, progress is hindered by lack of large, high-quality datasets and benchmarks. In addition to poor-quality due to challenges in data collection, existing datasets suffer from leakage and unrealistic base rates, leading to overly optimistic performance results. In this paper, we introduce PhreshPhish, a large-scale, high-quality dataset of phishing websites that addresses these limitations. Compared to existing public datasets, PhreshPhish is substantially larger and provides significantly higher quality, as measured by the estimated rate of invalid or mislabeled data points. Additionally, we propose a comprehensive suite of benchmark datasets specifically designed for realistic model evaluation by minimizing leakage, increasing task difficulty, enhancing dataset diversity, and adjustment of base rates more likely to be seen in the real world. We train and evaluate multiple solution approaches to provide baseline performance on the benchmark sets. We believe the availability of this dataset and benchmarks will enable realistic, standardized model comparison and foster further advances in phishing detection. The datasets and benchmarks are available on Hugging Face (https://huggingface.co/datasets/phreshphish/phreshphish).
format Preprint
id arxiv_https___arxiv_org_abs_2507_10854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark
Dalton, Thomas
Gowda, Hemanth
Rao, Girish
Pargi, Sachin
Khodabakhshi, Alireza Hadj
Rombs, Joseph
Jou, Stephan
Marwah, Manish
Cryptography and Security
Artificial Intelligence
Machine Learning
Phishing remains a pervasive and growing threat, inflicting heavy economic and reputational damage. While machine learning has been effective in real-time detection of phishing attacks, progress is hindered by lack of large, high-quality datasets and benchmarks. In addition to poor-quality due to challenges in data collection, existing datasets suffer from leakage and unrealistic base rates, leading to overly optimistic performance results. In this paper, we introduce PhreshPhish, a large-scale, high-quality dataset of phishing websites that addresses these limitations. Compared to existing public datasets, PhreshPhish is substantially larger and provides significantly higher quality, as measured by the estimated rate of invalid or mislabeled data points. Additionally, we propose a comprehensive suite of benchmark datasets specifically designed for realistic model evaluation by minimizing leakage, increasing task difficulty, enhancing dataset diversity, and adjustment of base rates more likely to be seen in the real world. We train and evaluate multiple solution approaches to provide baseline performance on the benchmark sets. We believe the availability of this dataset and benchmarks will enable realistic, standardized model comparison and foster further advances in phishing detection. The datasets and benchmarks are available on Hugging Face (https://huggingface.co/datasets/phreshphish/phreshphish).
title PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.10854