PhishSSL: Self-Supervised Contrastive Learning for Phishing Website Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Wenhao, Manickam, Selvakumar, Chong, Yung-Wey, Karuppayah, Shankar, Nanda, Priyadarsi, Li, Binyong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916994563964928
author Li, Wenhao
Manickam, Selvakumar
Chong, Yung-Wey
Karuppayah, Shankar
Nanda, Priyadarsi
Li, Binyong
author_facet Li, Wenhao
Manickam, Selvakumar
Chong, Yung-Wey
Karuppayah, Shankar
Nanda, Priyadarsi
Li, Binyong
contents Phishing websites remain a persistent cybersecurity threat by mimicking legitimate sites to steal sensitive user information. Existing machine learning-based detection methods often rely on supervised learning with labeled data, which not only incurs substantial annotation costs but also limits adaptability to novel attack patterns. To address these challenges, we propose PhishSSL, a self-supervised contrastive learning framework that eliminates the need for labeled phishing data during training. PhishSSL combines hybrid tabular augmentation with adaptive feature attention to produce semantically consistent views and emphasize discriminative attributes. We evaluate PhishSSL on three phishing datasets with distinct feature compositions. Across all datasets, PhishSSL consistently outperforms unsupervised and self-supervised baselines, while ablation studies confirm the contribution of each component. Moreover, PhishSSL maintains robust performance despite the diversity of feature sets, highlighting its strong generalization and transferability. These results demonstrate that PhishSSL offers a promising solution for phishing website detection, particularly effective against evolving threats in dynamic Web environments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05900
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhishSSL: Self-Supervised Contrastive Learning for Phishing Website Detection
Li, Wenhao
Manickam, Selvakumar
Chong, Yung-Wey
Karuppayah, Shankar
Nanda, Priyadarsi
Li, Binyong
Cryptography and Security
Phishing websites remain a persistent cybersecurity threat by mimicking legitimate sites to steal sensitive user information. Existing machine learning-based detection methods often rely on supervised learning with labeled data, which not only incurs substantial annotation costs but also limits adaptability to novel attack patterns. To address these challenges, we propose PhishSSL, a self-supervised contrastive learning framework that eliminates the need for labeled phishing data during training. PhishSSL combines hybrid tabular augmentation with adaptive feature attention to produce semantically consistent views and emphasize discriminative attributes. We evaluate PhishSSL on three phishing datasets with distinct feature compositions. Across all datasets, PhishSSL consistently outperforms unsupervised and self-supervised baselines, while ablation studies confirm the contribution of each component. Moreover, PhishSSL maintains robust performance despite the diversity of feature sets, highlighting its strong generalization and transferability. These results demonstrate that PhishSSL offers a promising solution for phishing website detection, particularly effective against evolving threats in dynamic Web environments.
title PhishSSL: Self-Supervised Contrastive Learning for Phishing Website Detection
topic Cryptography and Security
url https://arxiv.org/abs/2510.05900