mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xiaona, Brif, Constantin, Lourentzou, Ismini
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914313023782912
author Zhou, Xiaona
Brif, Constantin
Lourentzou, Ismini
author_facet Zhou, Xiaona
Brif, Constantin
Lourentzou, Ismini
contents Anomaly detection in multivariate time series is essential across domains such as healthcare, cybersecurity, and industrial monitoring, yet remains fundamentally challenging due to high-dimensional dependencies, the presence of cross-correlations between time-dependent variables, and the scarcity of labeled anomalies. We introduce mTSBench, the largest benchmark to date for multivariate time series anomaly detection and model selection, consisting of 344 labeled time series across 19 datasets from a wide range of application domains. We comprehensively evaluate 24 anomaly detectors, including the only two publicly available large language model-based methods for multivariate time series. Consistent with prior findings, we observe that no single detector dominates across datasets, motivating the need for effective model selection. We benchmark three recent model selection methods and find that even the strongest of them remain far from optimal. Our results highlight the outstanding need for robust, generalizable selection strategies. We open-source the benchmark at https://plan-lab.github.io/mtsbench to encourage future research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21550
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
Zhou, Xiaona
Brif, Constantin
Lourentzou, Ismini
Machine Learning
Artificial Intelligence
Anomaly detection in multivariate time series is essential across domains such as healthcare, cybersecurity, and industrial monitoring, yet remains fundamentally challenging due to high-dimensional dependencies, the presence of cross-correlations between time-dependent variables, and the scarcity of labeled anomalies. We introduce mTSBench, the largest benchmark to date for multivariate time series anomaly detection and model selection, consisting of 344 labeled time series across 19 datasets from a wide range of application domains. We comprehensively evaluate 24 anomaly detectors, including the only two publicly available large language model-based methods for multivariate time series. Consistent with prior findings, we observe that no single detector dominates across datasets, motivating the need for effective model selection. We benchmark three recent model selection methods and find that even the strongest of them remain far from optimal. Our results highlight the outstanding need for robust, generalizable selection strategies. We open-source the benchmark at https://plan-lab.github.io/mtsbench to encourage future research.
title mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.21550