Semi-Supervised Supply Chain Fraud Detection with Unsupervised Pre-Filtering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moradi, Fatemeh, Tarif, Mehran, Homaei, Mohammadhossein
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908483430907904
author Moradi, Fatemeh
Tarif, Mehran
Homaei, Mohammadhossein
author_facet Moradi, Fatemeh
Tarif, Mehran
Homaei, Mohammadhossein
contents Detecting fraud in modern supply chains is a growing challenge, driven by the complexity of global networks and the scarcity of labeled data. Traditional detection methods often struggle with class imbalance and limited supervision, reducing their effectiveness in real-world applications. This paper proposes a novel two-phase learning framework to address these challenges. In the first phase, the Isolation Forest algorithm performs unsupervised anomaly detection to identify potential fraud cases and reduce the volume of data requiring further analysis. In the second phase, a self-training Support Vector Machine (SVM) refines the predictions using both labeled and high-confidence pseudo-labeled samples, enabling robust semi-supervised learning. The proposed method is evaluated on the DataCo Smart Supply Chain Dataset, a comprehensive real-world supply chain dataset with fraud indicators. It achieves an F1-score of 0.817 while maintaining a false positive rate below 3.0%. These results demonstrate the effectiveness and efficiency of combining unsupervised pre-filtering with semi-supervised refinement for supply chain fraud detection under real-world constraints, though we acknowledge limitations regarding concept drift and the need for comparison with deep learning approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-Supervised Supply Chain Fraud Detection with Unsupervised Pre-Filtering
Moradi, Fatemeh
Tarif, Mehran
Homaei, Mohammadhossein
Machine Learning
Cryptography and Security
Detecting fraud in modern supply chains is a growing challenge, driven by the complexity of global networks and the scarcity of labeled data. Traditional detection methods often struggle with class imbalance and limited supervision, reducing their effectiveness in real-world applications. This paper proposes a novel two-phase learning framework to address these challenges. In the first phase, the Isolation Forest algorithm performs unsupervised anomaly detection to identify potential fraud cases and reduce the volume of data requiring further analysis. In the second phase, a self-training Support Vector Machine (SVM) refines the predictions using both labeled and high-confidence pseudo-labeled samples, enabling robust semi-supervised learning. The proposed method is evaluated on the DataCo Smart Supply Chain Dataset, a comprehensive real-world supply chain dataset with fraud indicators. It achieves an F1-score of 0.817 while maintaining a false positive rate below 3.0%. These results demonstrate the effectiveness and efficiency of combining unsupervised pre-filtering with semi-supervised refinement for supply chain fraud detection under real-world constraints, though we acknowledge limitations regarding concept drift and the need for comparison with deep learning approaches.
title Semi-Supervised Supply Chain Fraud Detection with Unsupervised Pre-Filtering
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2508.06574