ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yoon, Sanghyu, Kim, Dongmin, Yoon, Suhee, Sim, Ye Seul, Yoa, Seungdong, Cho, Hye-Seung, Lee, Soonyoung, Lee, Hankook, Lim, Woohyung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914615053516800
author Yoon, Sanghyu
Kim, Dongmin
Yoon, Suhee
Sim, Ye Seul
Yoa, Seungdong
Cho, Hye-Seung
Lee, Soonyoung
Lee, Hankook
Lim, Woohyung
author_facet Yoon, Sanghyu
Kim, Dongmin
Yoon, Suhee
Sim, Ye Seul
Yoa, Seungdong
Cho, Hye-Seung
Lee, Soonyoung
Lee, Hankook
Lim, Woohyung
contents In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing benchmarks provide only raw data points without semantic context, overlooking rich textual metadata such as feature descriptions and domain knowledge that experts rely on in practice. This limitation restricts research flexibility and prevents models from fully leveraging domain knowledge for detection. ReTabAD addresses this gap by restoring textual semantics to enable context-aware tabular AD research. We provide (1) 20 carefully curated tabular datasets enriched with structured textual metadata, together with implementations of state-of-the-art AD algorithms including classical, deep learning, and LLM-based approaches, and (2) a zero-shot LLM framework that leverages semantic context without task-specific training, establishing a strong baseline for future research. Furthermore, this work provides insights into the role and utility of textual metadata in AD through experiments and analysis. Results show that semantic context improves detection performance and enhances interpretability by supporting domain-aware reasoning. These findings establish ReTabAD as a benchmark for systematic exploration of context-aware AD.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
Yoon, Sanghyu
Kim, Dongmin
Yoon, Suhee
Sim, Ye Seul
Yoa, Seungdong
Cho, Hye-Seung
Lee, Soonyoung
Lee, Hankook
Lim, Woohyung
Artificial Intelligence
Machine Learning
In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing benchmarks provide only raw data points without semantic context, overlooking rich textual metadata such as feature descriptions and domain knowledge that experts rely on in practice. This limitation restricts research flexibility and prevents models from fully leveraging domain knowledge for detection. ReTabAD addresses this gap by restoring textual semantics to enable context-aware tabular AD research. We provide (1) 20 carefully curated tabular datasets enriched with structured textual metadata, together with implementations of state-of-the-art AD algorithms including classical, deep learning, and LLM-based approaches, and (2) a zero-shot LLM framework that leverages semantic context without task-specific training, establishing a strong baseline for future research. Furthermore, this work provides insights into the role and utility of textual metadata in AD through experiments and analysis. Results show that semantic context improves detection performance and enhances interpretability by supporting domain-aware reasoning. These findings establish ReTabAD as a benchmark for systematic exploration of context-aware AD.
title ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.02060