Learning from sanctioned government suppliers: A machine learning and network science approach to detecting fraud and corruption in Mexico

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Medina-Hernández, Martí, Kertész, Janos, Fazekas, Mihály
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908732564176896
author Medina-Hernández, Martí
Kertész, Janos
Fazekas, Mihály
author_facet Medina-Hernández, Martí
Kertész, Janos
Fazekas, Mihály
contents Detecting fraud and corruption in public procurement remains a major challenge for governments worldwide. Most research to-date builds on domain-knowledge-based corruption risk indicators of individual contract-level features and some also analyzes contracting network patterns. A critical barrier for supervised machine learning is the absence of confirmed non-corrupt, negative, examples, which makes conventional machine learning inappropriate for this task. Using publicly available data on federally funded procurement in Mexico and company sanction records, this study implements positive-unlabeled (PU) learning algorithms that integrate domain-knowledge-based red flags with network-derived features to identify likely corrupt and fraudulent contracts. The best-performing PU model on average captures 32 percent more known positives and performs on average 2.3 times better than random guessing, substantially outperforming approaches based solely on traditional red flags. The analysis of the Shapley Additive Explanations reveals that network-derived features, particularly those associated with contracts in the network core or suppliers with high eigenvector centrality, are the most important. Traditional red flags further enhance model performance in line with expectations, albeit mainly for contracts awarded through competitive tenders. This methodology can support law enforcement in Mexico, and it can be adapted to other national contexts too.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19491
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning from sanctioned government suppliers: A machine learning and network science approach to detecting fraud and corruption in Mexico
Medina-Hernández, Martí
Kertész, Janos
Fazekas, Mihály
Machine Learning
Computers and Society
Detecting fraud and corruption in public procurement remains a major challenge for governments worldwide. Most research to-date builds on domain-knowledge-based corruption risk indicators of individual contract-level features and some also analyzes contracting network patterns. A critical barrier for supervised machine learning is the absence of confirmed non-corrupt, negative, examples, which makes conventional machine learning inappropriate for this task. Using publicly available data on federally funded procurement in Mexico and company sanction records, this study implements positive-unlabeled (PU) learning algorithms that integrate domain-knowledge-based red flags with network-derived features to identify likely corrupt and fraudulent contracts. The best-performing PU model on average captures 32 percent more known positives and performs on average 2.3 times better than random guessing, substantially outperforming approaches based solely on traditional red flags. The analysis of the Shapley Additive Explanations reveals that network-derived features, particularly those associated with contracts in the network core or suppliers with high eigenvector centrality, are the most important. Traditional red flags further enhance model performance in line with expectations, albeit mainly for contracts awarded through competitive tenders. This methodology can support law enforcement in Mexico, and it can be adapted to other national contexts too.
title Learning from sanctioned government suppliers: A machine learning and network science approach to detecting fraud and corruption in Mexico
topic Machine Learning
Computers and Society
url https://arxiv.org/abs/2512.19491