Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | McKechnie, Jack, McDonald, Graham, Macdonald, Craig |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SARA: A Collection of Sensitivity-Aware Relevance Assessments
por: McKechnie, Jack, et al.
Publicado: (2024)
por: McKechnie, Jack, et al.
Publicado: (2024)
Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias
por: Dehghan, Mahdi, et al.
Publicado: (2026)
por: Dehghan, Mahdi, et al.
Publicado: (2026)
Document Similarity Enhanced IPS Estimation for Unbiased Learning to Rank
por: Liang, Zeyan, et al.
Publicado: (2025)
por: Liang, Zeyan, et al.
Publicado: (2025)
Query Exposure Prediction for Groups of Documents in Rankings
por: Jaenich, Thomas, et al.
Publicado: (2024)
por: Jaenich, Thomas, et al.
Publicado: (2024)
What Else Would I Like? A User Simulator using Alternatives for Improved Evaluation of Fashion Conversational Recommendation Systems
por: Vlachou, Maria, et al.
Publicado: (2024)
por: Vlachou, Maria, et al.
Publicado: (2024)
Predicting Retrieval Utility and Answer Quality in Retrieval-Augmented Generation
por: Tian, Fangzheng, et al.
Publicado: (2026)
por: Tian, Fangzheng, et al.
Publicado: (2026)
On Precomputation and Caching in Information Retrieval Experiments with Pipeline Architectures
por: MacAvaney, Sean, et al.
Publicado: (2025)
por: MacAvaney, Sean, et al.
Publicado: (2025)
RAQG-QPP: Query Performance Prediction with Retrieved Query Variants and Retrieval Augmented Query Generation
por: Tian, Fangzheng, et al.
Publicado: (2026)
por: Tian, Fangzheng, et al.
Publicado: (2026)
Is Relevance Propagated from Retriever to Generator in RAG?
por: Tian, Fangzheng, et al.
Publicado: (2025)
por: Tian, Fangzheng, et al.
Publicado: (2025)
Revisiting Query Variants: The Advantage of Retrieval Over Generation of Query Variants for Effective QPP
por: Tian, Fangzheng, et al.
Publicado: (2025)
por: Tian, Fangzheng, et al.
Publicado: (2025)
All Eyes on the Ranker: Participatory Auditing to Surface Blind Spots in Ranked Search Results
por: Rezk, Anna Marie, et al.
Publicado: (2026)
por: Rezk, Anna Marie, et al.
Publicado: (2026)
Temporal Fact Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN
por: Dey, Ritajit, et al.
Publicado: (2026)
por: Dey, Ritajit, et al.
Publicado: (2026)
Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
por: Macdonald, Craig, et al.
Publicado: (2025)
por: Macdonald, Craig, et al.
Publicado: (2025)
Shallow Cross-Encoders for Low-Latency Retrieval
por: Petrov, Aleksandr V., et al.
Publicado: (2024)
por: Petrov, Aleksandr V., et al.
Publicado: (2024)
Beyond Questions: Leveraging ColBERT for Keyphrase Search
por: Gabín, Jorge, et al.
Publicado: (2024)
por: Gabín, Jorge, et al.
Publicado: (2024)
Quantifying Query Fairness Under Unawareness
por: Jaenich, Thomas, et al.
Publicado: (2025)
por: Jaenich, Thomas, et al.
Publicado: (2025)
Aligning GPTRec with Beyond-Accuracy Goals with Reinforcement Learning
por: Petrov, Aleksandr, et al.
Publicado: (2024)
por: Petrov, Aleksandr, et al.
Publicado: (2024)
Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Embeddings
por: Petrov, Aleksandr V., et al.
Publicado: (2025)
por: Petrov, Aleksandr V., et al.
Publicado: (2025)
Pipeline Inspection, Visualization, and Interoperability in PyTerrier
por: Lionis, Emmanouil Georgios, et al.
Publicado: (2026)
por: Lionis, Emmanouil Georgios, et al.
Publicado: (2026)
Neural Passage Quality Estimation for Static Pruning
por: Chang, Xuejun, et al.
Publicado: (2024)
por: Chang, Xuejun, et al.
Publicado: (2024)
Replicability Measures for Longitudinal Information Retrieval Evaluation
por: Keller, Jüri, et al.
Publicado: (2024)
por: Keller, Jüri, et al.
Publicado: (2024)
Am I on the Right Track? What Can Predicted Query Performance Tell Us about the Search Behaviour of Agentic RAG
por: Tian, Fangzheng, et al.
Publicado: (2025)
por: Tian, Fangzheng, et al.
Publicado: (2025)
GenTREC: The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2025)
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2025)
Towards Reliable Testing for Multiple Information Retrieval System Comparisons
por: Otero, David, et al.
Publicado: (2025)
por: Otero, David, et al.
Publicado: (2025)
Synthetic Test Collections for Retrieval Evaluation
por: Rahmani, Hossein A., et al.
Publicado: (2024)
por: Rahmani, Hossein A., et al.
Publicado: (2024)
Efficient Inference of Sub-Item Id-based Sequential Recommendation Models with Millions of Items
por: Petrov, Aleksandr V., et al.
Publicado: (2024)
por: Petrov, Aleksandr V., et al.
Publicado: (2024)
SelRoute: Query-Type-Aware Routing for Long-Term Conversational Memory Retrieval
por: McKee, Matthew
Publicado: (2026)
por: McKee, Matthew
Publicado: (2026)
Brain-Machine Interfaces & Information Retrieval Challenges and Opportunities
por: Moshfeghi, Yashar, et al.
Publicado: (2025)
por: Moshfeghi, Yashar, et al.
Publicado: (2025)
Enhancing Sequential Music Recommendation with Personalized Popularity Awareness
por: Abbattista, Davide, et al.
Publicado: (2024)
por: Abbattista, Davide, et al.
Publicado: (2024)
Self-Service Circulation: An Exploratory Study.
por: Carey, Robert F., et al.
Publicado: (1998)
por: Carey, Robert F., et al.
Publicado: (1998)
Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
por: Tyagin, Ilya, et al.
Publicado: (2025)
por: Tyagin, Ilya, et al.
Publicado: (2025)
Towards Brain Passage Retrieval -- An Investigation of EEG Query Representations
por: McGuire, Niall, et al.
Publicado: (2024)
por: McGuire, Niall, et al.
Publicado: (2024)
Offline Evaluation Measures of Fairness in Recommender Systems
por: Rampisela, Theresia Veronika
Publicado: (2026)
por: Rampisela, Theresia Veronika
Publicado: (2026)
Dense Passage Retrieval in Conversational Search
por: Salamah, Ahmed H., et al.
Publicado: (2025)
por: Salamah, Ahmed H., et al.
Publicado: (2025)
A Systematic Study of Biomedical Retrieval Pipeline Trade-offs in Performance and Efficiency
por: Stepanyan, Hayk, et al.
Publicado: (2026)
por: Stepanyan, Hayk, et al.
Publicado: (2026)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
por: Arabzadeh, Negar, et al.
Publicado: (2024)
por: Arabzadeh, Negar, et al.
Publicado: (2024)
Retrieval of ERIC Files. An On-Line Approach.
por: McIsaac, Donald N., et al.
Publicado: (1973)
por: McIsaac, Donald N., et al.
Publicado: (1973)
NeuCLIRBench: A Modern Evaluation Collection for Monolingual, Cross-Language, and Multilingual Information Retrieval
por: Lawrie, Dawn, et al.
Publicado: (2025)
por: Lawrie, Dawn, et al.
Publicado: (2025)
NeuCLIRTech: Chinese Monolingual and Cross-Language Information Retrieval Evaluation in a Challenging Domain
por: Lawrie, Dawn, et al.
Publicado: (2026)
por: Lawrie, Dawn, et al.
Publicado: (2026)
Human-Computer Interaction as a basis for assessing Geographic Information Retrieval Systems.
por: Manuel Enrique Puebla Martínez
Publicado: (2018)
por: Manuel Enrique Puebla Martínez
Publicado: (2018)
Ejemplares similares
-
SARA: A Collection of Sensitivity-Aware Relevance Assessments
por: McKechnie, Jack, et al.
Publicado: (2024) -
Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias
por: Dehghan, Mahdi, et al.
Publicado: (2026) -
Document Similarity Enhanced IPS Estimation for Unbiased Learning to Rank
por: Liang, Zeyan, et al.
Publicado: (2025) -
Query Exposure Prediction for Groups of Documents in Rankings
por: Jaenich, Thomas, et al.
Publicado: (2024) -
What Else Would I Like? A User Simulator using Alternatives for Improved Evaluation of Fashion Conversational Recommendation Systems
por: Vlachou, Maria, et al.
Publicado: (2024)