A Fisher's exact test justification of the TF-IDF term-weighting scheme
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sheridan, Paul, Ahmed, Zeyad, Farooque, Aitazaz A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
von: Ahmed, Zeyad, et al.
Veröffentlicht: (2026)
von: Ahmed, Zeyad, et al.
Veröffentlicht: (2026)
Non-Random Data Encodes its Geometric and Topological Dimensions
von: Zenil, Hector, et al.
Veröffentlicht: (2024)
von: Zenil, Hector, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
von: Guo, Yang, et al.
Veröffentlicht: (2025)
von: Guo, Yang, et al.
Veröffentlicht: (2025)
A Unified Bayesian Perspective for Conventional and Robust Adaptive Filters
von: Szczecinski, Leszek, et al.
Veröffentlicht: (2025)
von: Szczecinski, Leszek, et al.
Veröffentlicht: (2025)
MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG
von: Wang, Xihang, et al.
Veröffentlicht: (2026)
von: Wang, Xihang, et al.
Veröffentlicht: (2026)
A statistical significance testing approach for measuring term burstiness with applications to domain-specific terminology extraction
von: Hurtado, Samuel Sarria, et al.
Veröffentlicht: (2023)
von: Hurtado, Samuel Sarria, et al.
Veröffentlicht: (2023)
Bradford's Law: Theory, Empiricism and the Gaps Between.
von: Drott, M. Carl
Veröffentlicht: (1981)
von: Drott, M. Carl
Veröffentlicht: (1981)
A quantum semantic framework for natural language processing
von: Agostino, Christopher J., et al.
Veröffentlicht: (2025)
von: Agostino, Christopher J., et al.
Veröffentlicht: (2025)
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
Mitigating Hallucination with ZeroG: An Advanced Knowledge Management Engine
von: Sharma, Anantha, et al.
Veröffentlicht: (2024)
von: Sharma, Anantha, et al.
Veröffentlicht: (2024)
Information-Theoretic Generative Clustering of Documents
von: Du, Xin, et al.
Veröffentlicht: (2024)
von: Du, Xin, et al.
Veröffentlicht: (2024)
Improving Robustness of Tabular Retrieval via Representational Stability
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2026)
von: Bhandari, Kushal Raj, et al.
Veröffentlicht: (2026)
HRGraph: Leveraging LLMs for HR Data Knowledge Graphs with Information Propagation-based Job Recommendation
von: Wasi, Azmine Toushik
Veröffentlicht: (2024)
von: Wasi, Azmine Toushik
Veröffentlicht: (2024)
LM4OPT: Unveiling the Potential of Large Language Models in Formulating Mathematical Optimization Problems
von: Ahmed, Tasnim, et al.
Veröffentlicht: (2024)
von: Ahmed, Tasnim, et al.
Veröffentlicht: (2024)
Enhancing Medical Support in the Arabic Language Through Personalized ChatGPT Assistance
von: Issa, Mohamed, et al.
Veröffentlicht: (2024)
von: Issa, Mohamed, et al.
Veröffentlicht: (2024)
FinBloom: Knowledge Grounding Large Language Model with Real-time Financial Data
von: Sinha, Ankur, et al.
Veröffentlicht: (2025)
von: Sinha, Ankur, et al.
Veröffentlicht: (2025)
Prompt-Based LLMs for Position Bias-Aware Reranking in Personalized Recommendations
von: Islam, Md Aminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Aminul, et al.
Veröffentlicht: (2025)
A Decision Theory View of the Information Retrieval Situation: An Operations Research Approach
von: Kraft, Donald H.
Veröffentlicht: (1973)
von: Kraft, Donald H.
Veröffentlicht: (1973)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
von: Levy, Jordan, et al.
Veröffentlicht: (2026)
von: Levy, Jordan, et al.
Veröffentlicht: (2026)
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
von: Flor, Michael, et al.
Veröffentlicht: (2025)
von: Flor, Michael, et al.
Veröffentlicht: (2025)
Disease Identification From Unstructured User Input
von: Faisal, Fahim, et al.
Veröffentlicht: (2019)
von: Faisal, Fahim, et al.
Veröffentlicht: (2019)
RAC: Retrieval-Augmented Clarification for Faithful Conversational Search
von: Kebir, Ahmed Rayane, et al.
Veröffentlicht: (2026)
von: Kebir, Ahmed Rayane, et al.
Veröffentlicht: (2026)
Signal in the Noise: Decoding the Reality of Airline Service Quality with Large Language Models
von: Dawoud, Ahmed, et al.
Veröffentlicht: (2026)
von: Dawoud, Ahmed, et al.
Veröffentlicht: (2026)
Cross-modal Retrieval for Knowledge-based Visual Question Answering
von: Lerner, Paul, et al.
Veröffentlicht: (2024)
von: Lerner, Paul, et al.
Veröffentlicht: (2024)
Exploring Traffic Crash Narratives in Jordan Using Text Mining Analytics
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
von: Jaradat, Shadi, et al.
Veröffentlicht: (2024)
Detection of Temporality at Discourse Level on Financial News by Combining Natural Language Processing and Machine Learning
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
Automatic detection of relevant information, predictions and forecasts in financial news through topic modelling with Latent Dirichlet Allocation
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
von: García-Méndez, Silvia, et al.
Veröffentlicht: (2024)
A comparison of latent semantic analysis and correspondence analysis of document-term matrices
von: Qi, Qianqian, et al.
Veröffentlicht: (2021)
von: Qi, Qianqian, et al.
Veröffentlicht: (2021)
Extending Translate-Train for ColBERT-X to African Language CLIR
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
von: Yang, Eugene, et al.
Veröffentlicht: (2024)
A smoothed-Bayesian approach to frequency recovery from sketched data
von: Beraha, Mario, et al.
Veröffentlicht: (2023)
von: Beraha, Mario, et al.
Veröffentlicht: (2023)
Stemming -- The Evolution and Current State with a Focus on Bangla
von: Paul, Abhijit, et al.
Veröffentlicht: (2025)
von: Paul, Abhijit, et al.
Veröffentlicht: (2025)
New Directions in Text Classification Research: Maximizing The Performance of Sentiment Classification from Limited Data
von: Agustian, Surya, et al.
Veröffentlicht: (2024)
von: Agustian, Surya, et al.
Veröffentlicht: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
Hybrid Student-Teacher Large Language Model Refinement for Cancer Toxicity Symptom Extraction
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2024)
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2024)
The Kolmogorov Complexity of Irish traditional dance music
von: McGettrick, Michael, et al.
Veröffentlicht: (2024)
von: McGettrick, Michael, et al.
Veröffentlicht: (2024)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
A Reproducibility Study of LLM-Based Query Reformulation
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies
von: Myers, Skatje, et al.
Veröffentlicht: (2024)
von: Myers, Skatje, et al.
Veröffentlicht: (2024)
On the Evaluation of Machine-Generated Reports
von: Mayfield, James, et al.
Veröffentlicht: (2024)
von: Mayfield, James, et al.
Veröffentlicht: (2024)
Back-of-the-Book Index Automation for Arabic Documents
von: Haidar, Nawal, et al.
Veröffentlicht: (2024)
von: Haidar, Nawal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Common TF-IDF variants arise as key components in the test statistic of a penalized likelihood-ratio test for word burstiness
von: Ahmed, Zeyad, et al.
Veröffentlicht: (2026) -
Non-Random Data Encodes its Geometric and Topological Dimensions
von: Zenil, Hector, et al.
Veröffentlicht: (2024) -
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
von: Guo, Yang, et al.
Veröffentlicht: (2025) -
A Unified Bayesian Perspective for Conventional and Robust Adaptive Filters
von: Szczecinski, Leszek, et al.
Veröffentlicht: (2025) -
MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG
von: Wang, Xihang, et al.
Veröffentlicht: (2026)