Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ahmed, Zeyad, Sheridan, Paul, McIsaac, Michael, Farooque, Aitazaz A.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911568370860032
author Ahmed, Zeyad
Sheridan, Paul
McIsaac, Michael
Farooque, Aitazaz A.
author_facet Ahmed, Zeyad
Sheridan, Paul
McIsaac, Michael
Farooque, Aitazaz A.
contents TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00672
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
Ahmed, Zeyad
Sheridan, Paul
McIsaac, Michael
Farooque, Aitazaz A.
Computation and Language
Information Retrieval
Statistics Theory
TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development.
title Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
topic Computation and Language
Information Retrieval
Statistics Theory
url https://arxiv.org/abs/2604.00672