Towards Token-Level Text Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Yang, Yu, Bicheng, Yang, Sikun, Liu, Ming, Yang, Yujiu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909995579211776
author Cao, Yang
Yu, Bicheng
Yang, Sikun
Liu, Ming
Yang, Yujiu
author_facet Cao, Yang
Yu, Bicheng
Yang, Sikun
Liu, Ming
Yang, Yujiu
contents Despite significant progress in text anomaly detection for web applications such as spam filtering and fake news detection, existing methods are fundamentally limited to document-level analysis, unable to identify which specific parts of a text are anomalous. We introduce token-level anomaly detection, a novel paradigm that enables fine-grained localization of anomalies within text. We formally define text anomalies at both document and token-levels, and propose a unified detection framework that operates across multiple levels. To facilitate research in this direction, we collect and annotate three benchmark datasets spanning spam, reviews and grammar errors with token-level labels. Experimental results demonstrate that our framework get better performance than other 6 baselines, opening new possibilities for precise anomaly localization in text. All the codes and data are publicly available on https://github.com/charles-cao/TokenCore.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13644
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Token-Level Text Anomaly Detection
Cao, Yang
Yu, Bicheng
Yang, Sikun
Liu, Ming
Yang, Yujiu
Computation and Language
Machine Learning
Despite significant progress in text anomaly detection for web applications such as spam filtering and fake news detection, existing methods are fundamentally limited to document-level analysis, unable to identify which specific parts of a text are anomalous. We introduce token-level anomaly detection, a novel paradigm that enables fine-grained localization of anomalies within text. We formally define text anomalies at both document and token-levels, and propose a unified detection framework that operates across multiple levels. To facilitate research in this direction, we collect and annotate three benchmark datasets spanning spam, reviews and grammar errors with token-level labels. Experimental results demonstrate that our framework get better performance than other 6 baselines, opening new possibilities for precise anomaly localization in text. All the codes and data are publicly available on https://github.com/charles-cao/TokenCore.
title Towards Token-Level Text Anomaly Detection
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.13644