JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929636228726784 |
|---|---|
| author | Rahmani, Hossein A. Yilmaz, Emine Craswell, Nick Mitra, Bhaskar |
| author_facet | Rahmani, Hossein A. Yilmaz, Emine Craswell, Nick Mitra, Bhaskar |
| contents | The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_13268 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment Rahmani, Hossein A. Yilmaz, Emine Craswell, Nick Mitra, Bhaskar Information Retrieval The effective training and evaluation of retrieval systems require a substantial amount of relevance judgments, which are traditionally collected from human assessors -- a process that is both costly and time-consuming. Large Language Models (LLMs) have shown promise in generating relevance labels for search tasks, offering a potential alternative to manual assessments. Current approaches often rely on a single LLM, such as GPT-4, which, despite being effective, are expensive and prone to intra-model biases that can favour systems leveraging similar models. In this work, we introduce JudgeBlender, a framework that employs smaller, open-source models to provide relevance judgments by combining evaluations across multiple LLMs (LLMBlender) or multiple prompts (PromptBlender). By leveraging the LLMJudge benchmark [18], we compare JudgeBlender with state-of-the-art methods and the top performers in the LLMJudge challenge. Our results show that JudgeBlender achieves competitive performance, demonstrating that very large models are often unnecessary for reliable relevance assessments. |
| title | JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2412.13268 |