Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
Fuente:
arXiv
Saved in:
| Main Authors: | Rahmani, Hossein A., Siro, Clemencia, Aliannejadi, Mohammad, Craswell, Nick, Clarke, Charles L. A., Faggioli, Guglielmo, Mitra, Bhaskar, Thomas, Paul, Yilmaz, Emine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Towards Understanding Bias in Synthetic Data for Evaluation
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
Overview of the TREC 2021 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Overview of the TREC 2023 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Towards Group-aware Search Success
by: Wu, Haolun, et al.
Published: (2024)
by: Wu, Haolun, et al.
Published: (2024)
Report on the Workshop on Simulations for Information Access (Sim4IA 2024) at SIGIR 2024
by: Breuer, Timo, et al.
Published: (2024)
by: Breuer, Timo, et al.
Published: (2024)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Overview of the TREC 2022 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Robust-IR @ SIGIR 2025: The First Workshop on Robust Information Retrieval
by: Liu, Yu-An, et al.
Published: (2025)
by: Liu, Yu-An, et al.
Published: (2025)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
by: Ramezan, Kimia, et al.
Published: (2025)
by: Ramezan, Kimia, et al.
Published: (2025)
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023)
by: Thomas, Paul, et al.
Published: (2023)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Words Blending Boxes. Obfuscating Queries in Information Retrieval using Differential Privacy
by: De Faveri, Francesco Luigi, et al.
Published: (2024)
by: De Faveri, Francesco Luigi, et al.
Published: (2024)
Emancipatory Information Retrieval
by: Mitra, Bhaskar
Published: (2025)
by: Mitra, Bhaskar
Published: (2025)
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
KEIR @ ECIR 2025: The Second Workshop on Knowledge-Enhanced Information Retrieval
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Understanding the Role of User Profile in the Personalization of Large Language Models
by: Wu, Bin, et al.
Published: (2024)
by: Wu, Bin, et al.
Published: (2024)
Interactions with Generative Information Retrieval Systems
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
IRLab@iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
by: Dietz, Laura, et al.
Published: (2025)
by: Dietz, Laura, et al.
Published: (2025)
Second SIGIR Workshop on Simulations for Information Access (Sim4IA 2025)
by: Schaer, Philipp, et al.
Published: (2025)
by: Schaer, Philipp, et al.
Published: (2025)
Investigating LLM Variability in Personalized Conversational Information Retrieval
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
Search and Society: Reimagining Information Access for Radical Futures
by: Mitra, Bhaskar
Published: (2024)
by: Mitra, Bhaskar
Published: (2024)
RMIT-ADM+S at the SIGIR 2025 LiveRAG Challenge
by: Ran, Kun, et al.
Published: (2025)
by: Ran, Kun, et al.
Published: (2025)
Controlling Gender Bias in Retrieval via a Backpack Architecture
by: Afzali, Amirabbas, et al.
Published: (2025)
by: Afzali, Amirabbas, et al.
Published: (2025)
Adaptive Retrieval-Augmented Generation for Conversational Systems
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
by: Qiao, Shuofei, et al.
Published: (2026)
by: Qiao, Shuofei, et al.
Published: (2026)
Similar Items
-
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025) -
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024) -
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024) -
Synthetic Test Collections for Retrieval Evaluation
by: Rahmani, Hossein A., et al.
Published: (2024) -
SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval
by: Rahmani, Hossein A., et al.
Published: (2024)