Saved in:
Bibliographic Details
Main Authors: Marwah, Manish, Narayanan, Asad, Jou, Stephan, Arlitt, Martin, Pospelova, Maria
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.14664
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910547955417088
author Marwah, Manish
Narayanan, Asad
Jou, Stephan
Arlitt, Martin
Pospelova, Maria
author_facet Marwah, Manish
Narayanan, Asad
Jou, Stephan
Arlitt, Martin
Pospelova, Maria
contents The cost of errors related to machine learning classifiers, namely, false positives and false negatives, are not equal and are application dependent. For example, in cybersecurity applications, the cost of not detecting an attack is very different from marking a benign activity as an attack. Various design choices during machine learning model building, such as hyperparameter tuning and model selection, allow a data scientist to trade-off between these two errors. However, most of the commonly used metrics to evaluate model quality, such as $F_1$ score, which is defined in terms of model precision and recall, treat both these errors equally, making it difficult for users to optimize for the actual cost of these errors. In this paper, we propose a new cost-aware metric, $C_{score}$ based on precision and recall that can replace $F_1$ score for model evaluation and selection. It includes a cost ratio that takes into account the differing costs of handling false positives and false negatives. We derive and characterize the new cost metric, and compare it to $F_1$ score. Further, we use this metric for model thresholding for five cybersecurity related datasets for multiple cost ratios. The results show an average cost savings of 49%.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14664
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Is $F_1$ Score Suboptimal for Cybersecurity Models? Introducing $C_{score}$, a Cost-Aware Alternative for Model Assessment
Marwah, Manish
Narayanan, Asad
Jou, Stephan
Arlitt, Martin
Pospelova, Maria
Machine Learning
Artificial Intelligence
The cost of errors related to machine learning classifiers, namely, false positives and false negatives, are not equal and are application dependent. For example, in cybersecurity applications, the cost of not detecting an attack is very different from marking a benign activity as an attack. Various design choices during machine learning model building, such as hyperparameter tuning and model selection, allow a data scientist to trade-off between these two errors. However, most of the commonly used metrics to evaluate model quality, such as $F_1$ score, which is defined in terms of model precision and recall, treat both these errors equally, making it difficult for users to optimize for the actual cost of these errors. In this paper, we propose a new cost-aware metric, $C_{score}$ based on precision and recall that can replace $F_1$ score for model evaluation and selection. It includes a cost ratio that takes into account the differing costs of handling false positives and false negatives. We derive and characterize the new cost metric, and compare it to $F_1$ score. Further, we use this metric for model thresholding for five cybersecurity related datasets for multiple cost ratios. The results show an average cost savings of 49%.
title Is $F_1$ Score Suboptimal for Cybersecurity Models? Introducing $C_{score}$, a Cost-Aware Alternative for Model Assessment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2407.14664