Tide: A Customisable Dataset Generator for Anti-Money Laundering Research

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Beukel, Montijn van den, Rožanec, Jože Martin, Varbanescu, Ana-Lucia
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910037937487872
author Beukel, Montijn van den
Rožanec, Jože Martin
Varbanescu, Ana-Lucia
author_facet Beukel, Montijn van den
Rožanec, Jože Martin
Varbanescu, Ana-Lucia
contents The lack of accessible transactional data significantly hinders machine learning research for Anti-Money Laundering (AML). Privacy and legal concerns prevent the sharing of real financial data, while existing synthetic generators focus on simplistic structural patterns and neglect the temporal dynamics (timing and frequency) that characterise sophisticated laundering schemes. We present Tide, an open-source synthetic dataset generator that produces graph-based financial networks incorporating money laundering patterns defined by both structural and temporal characteristics. Tide enables reproducible, customisable dataset generation tailored to specific research needs. We release two reference datasets with varying illicit ratios (LI: 0.10\%, HI: 0.19\%), alongside the implementation of state-of-the-art detection models. Evaluation across these datasets reveals condition-dependent model rankings: LightGBM achieves the highest PR-AUC (78.05) in the low illicit ratio condition, while XGBoost performs best (85.12) at higher fraud prevalence. These divergent rankings demonstrate that the reference datasets can meaningfully differentiate model capabilities across operational conditions. Tide provides the research community with a configurable benchmark that exposes meaningful performance variation across model architectures, advancing the development of robust AML detection methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tide: A Customisable Dataset Generator for Anti-Money Laundering Research
Beukel, Montijn van den
Rožanec, Jože Martin
Varbanescu, Ana-Lucia
Machine Learning
Artificial Intelligence
The lack of accessible transactional data significantly hinders machine learning research for Anti-Money Laundering (AML). Privacy and legal concerns prevent the sharing of real financial data, while existing synthetic generators focus on simplistic structural patterns and neglect the temporal dynamics (timing and frequency) that characterise sophisticated laundering schemes. We present Tide, an open-source synthetic dataset generator that produces graph-based financial networks incorporating money laundering patterns defined by both structural and temporal characteristics. Tide enables reproducible, customisable dataset generation tailored to specific research needs. We release two reference datasets with varying illicit ratios (LI: 0.10\%, HI: 0.19\%), alongside the implementation of state-of-the-art detection models. Evaluation across these datasets reveals condition-dependent model rankings: LightGBM achieves the highest PR-AUC (78.05) in the low illicit ratio condition, while XGBoost performs best (85.12) at higher fraud prevalence. These divergent rankings demonstrate that the reference datasets can meaningfully differentiate model capabilities across operational conditions. Tide provides the research community with a configurable benchmark that exposes meaningful performance variation across model architectures, advancing the development of robust AML detection methods.
title Tide: A Customisable Dataset Generator for Anti-Money Laundering Research
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.01863