Outlier Ranking in Large-Scale Public Health Streams

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Joshi, Ananya, Townes, Tina, Gormley, Nolan, Neureiter, Luke, Rosenfeld, Roni, Wilder, Bryan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909060081647616
author Joshi, Ananya
Townes, Tina
Gormley, Nolan
Neureiter, Luke
Rosenfeld, Roni
Wilder, Bryan
author_facet Joshi, Ananya
Townes, Tina
Gormley, Nolan
Neureiter, Luke
Rosenfeld, Roni
Wilder, Bryan
contents Disease control experts inspect public health data streams daily for outliers worth investigating, like those corresponding to data quality issues or disease outbreaks. However, they can only examine a few of the thousands of maximally-tied outliers returned by univariate outlier detection methods applied to large-scale public health data streams. To help experts distinguish the most important outliers from these thousands of tied outliers, we propose a new task for algorithms to rank the outputs of any univariate method applied to each of many streams. Our novel algorithm for this task, which leverages hierarchical networks and extreme value analysis, performed the best across traditional outlier detection metrics in a human-expert evaluation using public health data streams. Most importantly, experts have used our open-source Python implementation since April 2023 and report identifying outliers worth investigating 9.1x faster than their prior baseline. Other organizations can readily adapt this implementation to create rankings from the outputs of their tailored univariate methods across large-scale streams.
format Preprint
id arxiv_https___arxiv_org_abs_2401_01459
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Outlier Ranking in Large-Scale Public Health Streams
Joshi, Ananya
Townes, Tina
Gormley, Nolan
Neureiter, Luke
Rosenfeld, Roni
Wilder, Bryan
Artificial Intelligence
Disease control experts inspect public health data streams daily for outliers worth investigating, like those corresponding to data quality issues or disease outbreaks. However, they can only examine a few of the thousands of maximally-tied outliers returned by univariate outlier detection methods applied to large-scale public health data streams. To help experts distinguish the most important outliers from these thousands of tied outliers, we propose a new task for algorithms to rank the outputs of any univariate method applied to each of many streams. Our novel algorithm for this task, which leverages hierarchical networks and extreme value analysis, performed the best across traditional outlier detection metrics in a human-expert evaluation using public health data streams. Most importantly, experts have used our open-source Python implementation since April 2023 and report identifying outliers worth investigating 9.1x faster than their prior baseline. Other organizations can readily adapt this implementation to create rankings from the outputs of their tailored univariate methods across large-scale streams.
title Outlier Ranking in Large-Scale Public Health Streams
topic Artificial Intelligence
url https://arxiv.org/abs/2401.01459