Oddballness: universal anomaly detection with language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Graliński, Filip, Staruch, Ryszard, Jurkiewicz, Krzysztof
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916382802706432
author Graliński, Filip
Staruch, Ryszard
Jurkiewicz, Krzysztof
author_facet Graliński, Filip
Staruch, Ryszard
Jurkiewicz, Krzysztof
contents We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new metric introduced in this paper: oddballness. Oddballness measures how ``strange'' a given token is according to the language model. We demonstrate in grammatical error detection tasks (a specific case of text anomaly detection) that oddballness is better than just considering low-likelihood events, if a totally unsupervised setup is assumed.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03046
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Oddballness: universal anomaly detection with language models
Graliński, Filip
Staruch, Ryszard
Jurkiewicz, Krzysztof
Computation and Language
We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new metric introduced in this paper: oddballness. Oddballness measures how ``strange'' a given token is according to the language model. We demonstrate in grammatical error detection tasks (a specific case of text anomaly detection) that oddballness is better than just considering low-likelihood events, if a totally unsupervised setup is assumed.
title Oddballness: universal anomaly detection with language models
topic Computation and Language
url https://arxiv.org/abs/2409.03046