The Viability of Crowdsourcing for RAG Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Gienapp, Lukas, Hagen, Tim, Fröbe, Maik, Hagen, Matthias, Stein, Benno, Potthast, Martin, Scells, Harrisen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)
by: Gienapp, Lukas, et al.
Published: (2023)
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
by: Gienapp, Lukas, et al.
Published: (2024)
by: Gienapp, Lukas, et al.
Published: (2024)
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
by: Schlatt, Ferdinand, et al.
Published: (2024)
by: Schlatt, Ferdinand, et al.
Published: (2024)
Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders
by: Schlatt, Ferdinand, et al.
Published: (2024)
by: Schlatt, Ferdinand, et al.
Published: (2024)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
by: Parry, Andrew, et al.
Published: (2024)
by: Parry, Andrew, et al.
Published: (2024)
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
by: Schlatt, Ferdinand, et al.
Published: (2024)
by: Schlatt, Ferdinand, et al.
Published: (2024)
Investigating the Effects of Sparse Attention on Cross-Encoders
by: Schlatt, Ferdinand, et al.
Published: (2023)
by: Schlatt, Ferdinand, et al.
Published: (2023)
Detecting RAG Advertisements Across Advertising Styles
by: Heineking, Sebastian, et al.
Published: (2026)
by: Heineking, Sebastian, et al.
Published: (2026)
Counterfactual Query Rewriting to Use Historical Relevance Feedback
by: Keller, Jüri, et al.
Published: (2025)
by: Keller, Jüri, et al.
Published: (2025)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
by: Achkar, Pierre, et al.
Published: (2026)
by: Achkar, Pierre, et al.
Published: (2026)
AiReview: An Open Platform for Accelerating Systematic Reviews with LLMs
by: Mao, Xinyu, et al.
Published: (2025)
by: Mao, Xinyu, et al.
Published: (2025)
Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions
by: Heineking, Sebastian, et al.
Published: (2024)
by: Heineking, Sebastian, et al.
Published: (2024)
Detecting Generated Native Ads in Conversational Search
by: Schmidt, Sebastian, et al.
Published: (2024)
by: Schmidt, Sebastian, et al.
Published: (2024)
Variations in Relevance Judgments and the Shelf Life of Test Collections
by: Parry, Andrew, et al.
Published: (2025)
by: Parry, Andrew, et al.
Published: (2025)
Zero-shot Generative Large Language Models for Systematic Review Screening Automation
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
by: Polyakov, Gregory, et al.
Published: (2026)
by: Polyakov, Gregory, et al.
Published: (2026)
Simplified Longitudinal Retrieval Experiments: A Case Study on Query Expansion and Document Boosting
by: Keller, Jüri, et al.
Published: (2025)
by: Keller, Jüri, et al.
Published: (2025)
DenseReviewer: A Screening Prioritisation Tool for Systematic Review based on Dense Retrieval
by: Mao, Xinyu, et al.
Published: (2025)
by: Mao, Xinyu, et al.
Published: (2025)
AutoBool: An Reinforcement-Learning trained LLM for Effective Automated Boolean Query Generation for Systematic Reviews
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Overview of the Plagiarism Detection Task at PAN 2025
by: Greiner-Petter, André, et al.
Published: (2025)
by: Greiner-Petter, André, et al.
Published: (2025)
Investigating Counterclaims in Causality Extraction from Text
by: Hagen, Tim, et al.
Published: (2025)
by: Hagen, Tim, et al.
Published: (2025)
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
by: Rottach, Florian, et al.
Published: (2025)
by: Rottach, Florian, et al.
Published: (2025)
Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
by: He, Xuhong, et al.
Published: (2026)
by: He, Xuhong, et al.
Published: (2026)
A Collection of Systematic Reviews in Computer Science
by: Achkar, Pierre, et al.
Published: (2026)
by: Achkar, Pierre, et al.
Published: (2026)
Evaluating Search System Explainability with Psychometrics and Crowdsourcing
by: Chen, Catherine, et al.
Published: (2022)
by: Chen, Catherine, et al.
Published: (2022)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
by: Keller, Jüri, et al.
Published: (2026)
by: Keller, Jüri, et al.
Published: (2026)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Systems on Web and Scientific Data
by: Cancellieri, Matteo, et al.
Published: (2025)
by: Cancellieri, Matteo, et al.
Published: (2025)
Task-Oriented Paraphrase Analytics
by: Gohsen, Marcel, et al.
Published: (2024)
by: Gohsen, Marcel, et al.
Published: (2024)
RAG vs. GraphRAG: A Systematic Evaluation and Key Insights
by: Han, Haoyu, et al.
Published: (2025)
by: Han, Haoyu, et al.
Published: (2025)
Unlocking Crowdsourcing for Ontology Matching Validation
by: Qiang, Zhangcheng
Published: (2026)
by: Qiang, Zhangcheng
Published: (2026)
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
by: Cofala, Tim, et al.
Published: (2025)
by: Cofala, Tim, et al.
Published: (2025)
CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems
by: Bernard, Nolwenn, et al.
Published: (2024)
by: Bernard, Nolwenn, et al.
Published: (2024)
Constructing and Evaluating Declarative RAG Pipelines in PyTerrier
by: Macdonald, Craig, et al.
Published: (2025)
by: Macdonald, Craig, et al.
Published: (2025)
Controlled Retrieval-augmented Context Evaluation for Long-form RAG
by: Ju, Jia-Huei, et al.
Published: (2025)
by: Ju, Jia-Huei, et al.
Published: (2025)
Sensitivity-Aware Retrieval-Augmented Intent Clarification
by: Larooij, Maik
Published: (2026)
by: Larooij, Maik
Published: (2026)
LURE-RAG: Lightweight Utility-driven Reranking for Efficient RAG
by: Chandra, Manish, et al.
Published: (2026)
by: Chandra, Manish, et al.
Published: (2026)
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems
by: de Lima, Rafael Teixeira, et al.
Published: (2024)
by: de Lima, Rafael Teixeira, et al.
Published: (2024)
Similar Items
-
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023) -
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
by: Gienapp, Lukas, et al.
Published: (2024) -
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
by: Schlatt, Ferdinand, et al.
Published: (2024) -
Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders
by: Schlatt, Ferdinand, et al.
Published: (2024) -
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
by: Gienapp, Lukas, et al.
Published: (2025)