Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911568370860032 |
|---|---|
| author | Ahmed, Zeyad Sheridan, Paul McIsaac, Michael Farooque, Aitazaz A. |
| author_facet | Ahmed, Zeyad Sheridan, Paul McIsaac, Michael Farooque, Aitazaz A. |
| contents | TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_00672 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness Ahmed, Zeyad Sheridan, Paul McIsaac, Michael Farooque, Aitazaz A. Computation and Language Information Retrieval Statistics Theory TF-IDF is a classical formula that is widely used for identifying important terms within documents. We show that TF-IDF-like scores arise naturally from the test statistic of a penalized likelihood-ratio test setup capturing word burstiness (also known as word over-dispersion). In our framework, the alternative hypothesis captures word burstiness by modeling a collection of documents according to a family of beta-binomial distributions with a gamma penalty term on the precision parameter. In contrast, the null hypothesis assumes that words are binomially distributed in collection documents, a modeling approach that fails to account for word burstiness. We find that a term-weighting scheme given rise to by this test statistic performs comparably to TF-IDF on document classification tasks. This paper provides insights into TF-IDF from a statistical perspective and underscores the potential of hypothesis testing frameworks for advancing term-weighting scheme development. |
| title | Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness |
| topic | Computation and Language Information Retrieval Statistics Theory |
| url | https://arxiv.org/abs/2604.00672 |