Evaluation of Temporal Change in IR Test Collections
Fuente:
arXiv
Guardado en:
| Autores principales: | Keller, Jüri, Breuer, Timo, Schaer, Philipp |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Replicability Measures for Longitudinal Information Retrieval Evaluation
por: Keller, Jüri, et al.
Publicado: (2024)
por: Keller, Jüri, et al.
Publicado: (2024)
Evaluating Contrastive Feedback for Effective User Simulations
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025)
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
por: Keller, Jüri, et al.
Publicado: (2026)
por: Keller, Jüri, et al.
Publicado: (2026)
Context-Driven Interactive Query Simulations Based on Generative Large Language Models
por: Engelmann, Björn, et al.
Publicado: (2023)
por: Engelmann, Björn, et al.
Publicado: (2023)
Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025)
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025)
Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
por: Keller, Jüri, et al.
Publicado: (2025)
por: Keller, Jüri, et al.
Publicado: (2025)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
por: Cancellieri, Matteo, et al.
Publicado: (2025)
por: Cancellieri, Matteo, et al.
Publicado: (2025)
Counterfactual Query Rewriting to Use Historical Relevance Feedback
por: Keller, Jüri, et al.
Publicado: (2025)
por: Keller, Jüri, et al.
Publicado: (2025)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Systems on Web and Scientific Data
por: Cancellieri, Matteo, et al.
Publicado: (2025)
por: Cancellieri, Matteo, et al.
Publicado: (2025)
Data Fusion of Synthetic Query Variants With Generative Large Language Models
por: Breuer, Timo
Publicado: (2024)
por: Breuer, Timo
Publicado: (2024)
REANIMATOR: Reanimate Retrieval Test Collections with Extracted and Synthetic Resources
por: Engelmann, Björn, et al.
Publicado: (2025)
por: Engelmann, Björn, et al.
Publicado: (2025)
Second SIGIR Workshop on Simulations for Information Access (Sim4IA 2025)
por: Schaer, Philipp, et al.
Publicado: (2025)
por: Schaer, Philipp, et al.
Publicado: (2025)
Scientometric Analysis of the German IR Community within TREC & CLEF
por: Kruff, A. K., et al.
Publicado: (2025)
por: Kruff, A. K., et al.
Publicado: (2025)
Report on the Workshop on Simulations for Information Access (Sim4IA 2024) at SIGIR 2024
por: Breuer, Timo, et al.
Publicado: (2024)
por: Breuer, Timo, et al.
Publicado: (2024)
Validating Search Query Simulations: A Taxonomy of Measures
por: Kruff, Andreas Konstantin, et al.
Publicado: (2026)
por: Kruff, Andreas Konstantin, et al.
Publicado: (2026)
Auditing Search Query Suggestion Bias Through Recursive Algorithm Interrogation
por: Haak, Fabian, et al.
Publicado: (2026)
por: Haak, Fabian, et al.
Publicado: (2026)
LISP -- A Rich Interaction Dataset and Loggable Interactive Search Platform
por: Friese, Jana Isabelle, et al.
Publicado: (2026)
por: Friese, Jana Isabelle, et al.
Publicado: (2026)
Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR
por: Datta, Suchana, et al.
Publicado: (2026)
por: Datta, Suchana, et al.
Publicado: (2026)
SPECTRA: Synthetic IR Test Collections with Relevance Oracles and Controlled Distractor Diagnostics
por: Liang, Eric
Publicado: (2026)
por: Liang, Eric
Publicado: (2026)
Synthetic Test Collections for Retrieval Evaluation
por: Rahmani, Hossein A., et al.
Publicado: (2024)
por: Rahmani, Hossein A., et al.
Publicado: (2024)
ASPIRE: Assistive System for Performance Evaluation in IR
por: Peikos, Georgios, et al.
Publicado: (2024)
por: Peikos, Georgios, et al.
Publicado: (2024)
A Comparison of Methods for Evaluating Generative IR
por: Arabzadeh, Negar, et al.
Publicado: (2024)
por: Arabzadeh, Negar, et al.
Publicado: (2024)
LLM-Driven Usefulness Labeling for IR Evaluation
por: Dewan, Mouly, et al.
Publicado: (2025)
por: Dewan, Mouly, et al.
Publicado: (2025)
Ranking Narrative Query Graphs for Biomedical Document Retrieval (Technical Report)
por: Kroll, Hermann, et al.
Publicado: (2024)
por: Kroll, Hermann, et al.
Publicado: (2024)
GenTREC: The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2025)
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2025)
Building an Explainable Graph-based Biomedical Paper Recommendation System (Technical Report)
por: Kroll, Hermann, et al.
Publicado: (2024)
por: Kroll, Hermann, et al.
Publicado: (2024)
Improving the Reusability of Conversational Search Test Collections
por: Abbasiantaeb, Zahra, et al.
Publicado: (2025)
por: Abbasiantaeb, Zahra, et al.
Publicado: (2025)
Variations in Relevance Judgments and the Shelf Life of Test Collections
por: Parry, Andrew, et al.
Publicado: (2025)
por: Parry, Andrew, et al.
Publicado: (2025)
Investigating Bias in Political Search Query Suggestions by Relative Comparison with LLMs
por: Haak, Fabian, et al.
Publicado: (2024)
por: Haak, Fabian, et al.
Publicado: (2024)
DSEBench: A Test Collection for Explainable Dataset Search with Examples
por: Shi, Qing, et al.
Publicado: (2025)
por: Shi, Qing, et al.
Publicado: (2025)
Pairwise Comparison for Bias Identification and Quantification
por: Haak, Fabian, et al.
Publicado: (2025)
por: Haak, Fabian, et al.
Publicado: (2025)
STCALIR: Semi-Synthetic Test Collection for Algerian Legal Information Retrieval
por: Hatem, M'hamed Amine, et al.
Publicado: (2026)
por: Hatem, M'hamed Amine, et al.
Publicado: (2026)
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
por: Rahmani, Hossein A., et al.
Publicado: (2024)
por: Rahmani, Hossein A., et al.
Publicado: (2024)
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
por: Alaofi, Marwah, et al.
Publicado: (2025)
por: Alaofi, Marwah, et al.
Publicado: (2025)
SimEval-IR: A Unified Toolkit and Benchmark Suite for Evaluating User Simulators and Search Sessions
por: Zerhoudi, Saber
Publicado: (2026)
por: Zerhoudi, Saber
Publicado: (2026)
Identifying Offline Metrics that Predict Online Impact: A Pragmatic Strategy for Real-World Recommender Systems
por: Wilm, Timo, et al.
Publicado: (2025)
por: Wilm, Timo, et al.
Publicado: (2025)
Make Any Collection Navigable: Methods for Constructing and Evaluating Hypergraph of Text
por: Alvarez, Dean E., et al.
Publicado: (2026)
por: Alvarez, Dean E., et al.
Publicado: (2026)
ExcluIR: Exclusionary Neural Information Retrieval
por: Zhang, Wenhao, et al.
Publicado: (2024)
por: Zhang, Wenhao, et al.
Publicado: (2024)
Reproducing Adaptive Reranking for Reasoning-Intensive IR
por: Rathee, Mandeep, et al.
Publicado: (2026)
por: Rathee, Mandeep, et al.
Publicado: (2026)
Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
por: McKechnie, Jack, et al.
Publicado: (2025)
por: McKechnie, Jack, et al.
Publicado: (2025)
Ejemplares similares
-
Replicability Measures for Longitudinal Information Retrieval Evaluation
por: Keller, Jüri, et al.
Publicado: (2024) -
Evaluating Contrastive Feedback for Effective User Simulations
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025) -
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
por: Keller, Jüri, et al.
Publicado: (2026) -
Context-Driven Interactive Query Simulations Based on Generative Large Language Models
por: Engelmann, Björn, et al.
Publicado: (2023) -
Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
por: Kruff, Andreas Konstantin, et al.
Publicado: (2025)