AutoML in Cybersecurity: An Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saad, Sherif, Shi, Kevin, Mamun, Mohammed, Elmiligi, Hythem
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914061546946560
author Saad, Sherif
Shi, Kevin
Mamun, Mohammed
Elmiligi, Hythem
author_facet Saad, Sherif
Shi, Kevin
Mamun, Mohammed
Elmiligi, Hythem
contents Automated machine learning (AutoML) has emerged as a promising paradigm for automating machine learning (ML) pipeline design, broadening AI adoption. Yet its reliability in complex domains such as cybersecurity remains underexplored. This paper systematically evaluates eight open-source AutoML frameworks across 11 publicly available cybersecurity datasets, spanning intrusion detection, malware classification, phishing, fraud detection, and spam filtering. Results show substantial performance variability across tools and datasets, with no single solution consistently superior. A paradigm shift is observed: the challenge has moved from selecting individual ML models to identifying the most suitable AutoML framework, complicated by differences in runtime efficiency, automation capabilities, and supported features. AutoML tools frequently favor tree-based models, which perform well but risk overfitting and limit interpretability. Key challenges identified include adversarial vulnerability, model drift, and inadequate feature engineering. We conclude with best practices and research directions to strengthen robustness, interpretability, and trust in AutoML for high-stakes cybersecurity applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoML in Cybersecurity: An Empirical Study
Saad, Sherif
Shi, Kevin
Mamun, Mohammed
Elmiligi, Hythem
Cryptography and Security
Automated machine learning (AutoML) has emerged as a promising paradigm for automating machine learning (ML) pipeline design, broadening AI adoption. Yet its reliability in complex domains such as cybersecurity remains underexplored. This paper systematically evaluates eight open-source AutoML frameworks across 11 publicly available cybersecurity datasets, spanning intrusion detection, malware classification, phishing, fraud detection, and spam filtering. Results show substantial performance variability across tools and datasets, with no single solution consistently superior. A paradigm shift is observed: the challenge has moved from selecting individual ML models to identifying the most suitable AutoML framework, complicated by differences in runtime efficiency, automation capabilities, and supported features. AutoML tools frequently favor tree-based models, which perform well but risk overfitting and limit interpretability. Key challenges identified include adversarial vulnerability, model drift, and inadequate feature engineering. We conclude with best practices and research directions to strengthen robustness, interpretability, and trust in AutoML for high-stakes cybersecurity applications.
title AutoML in Cybersecurity: An Empirical Study
topic Cryptography and Security
url https://arxiv.org/abs/2509.23621