JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Rahmani, Hossein A., Yilmaz, Emine, Craswell, Nick, Mitra, Bhaskar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Towards Understanding Bias in Synthetic Data for Evaluation
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Overview of the TREC 2021 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Towards Group-aware Search Success
by: Wu, Haolun, et al.
Published: (2024)
by: Wu, Haolun, et al.
Published: (2024)
Overview of the TREC 2023 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Overview of the TREC 2022 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023)
by: Thomas, Paul, et al.
Published: (2023)
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Understanding the Role of User Profile in the Personalization of Large Language Models
by: Wu, Bin, et al.
Published: (2024)
by: Wu, Bin, et al.
Published: (2024)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Search and Society: Reimagining Information Access for Radical Futures
by: Mitra, Bhaskar
Published: (2024)
by: Mitra, Bhaskar
Published: (2024)
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
by: Yu, Chuting, et al.
Published: (2026)
by: Yu, Chuting, et al.
Published: (2026)
Don't Use LLMs to Make Relevance Judgments
by: Soboroff, Ian
Published: (2024)
by: Soboroff, Ian
Published: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)
by: de Jesus, Gabriel, et al.
Published: (2024)
Variations in Relevance Judgments and the Shelf Life of Test Collections
by: Parry, Andrew, et al.
Published: (2025)
by: Parry, Andrew, et al.
Published: (2025)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
A Study of Relevance Judgments.
by: Cuadra, Carlos A.
Published: (1968)
by: Cuadra, Carlos A.
Published: (1968)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Emancipatory Information Retrieval
by: Mitra, Bhaskar
Published: (2025)
by: Mitra, Bhaskar
Published: (2025)
Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation
by: Otero, David, et al.
Published: (2024)
by: Otero, David, et al.
Published: (2024)
TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
by: Chen, Nuo, et al.
Published: (2026)
by: Chen, Nuo, et al.
Published: (2026)
Open, to What End? A Capability-Theoretic Perspective on Open Search
by: Neophytou, Nicola, et al.
Published: (2026)
by: Neophytou, Nicola, et al.
Published: (2026)
Recall, Robustness, and Lexicographic Evaluation
by: Diaz, Fernando, et al.
Published: (2023)
by: Diaz, Fernando, et al.
Published: (2023)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
by: Ma, Shengjie, et al.
Published: (2024)
by: Ma, Shengjie, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Implicit Humanization in Everyday LLM Moral Judgments
by: Ayad, Hoda, et al.
Published: (2026)
by: Ayad, Hoda, et al.
Published: (2026)
Similar Items
-
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024) -
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025) -
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024) -
Towards Understanding Bias in Synthetic Data for Evaluation
by: Rahmani, Hossein A., et al.
Published: (2025) -
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
by: Rahmani, Hossein A., et al.
Published: (2024)