Can We Use Large Language Models to Fill Relevance Judgment Holes?
Fuente:
arXiv
Saved in:
| Main Authors: | Abbasiantaeb, Zahra, Meng, Chuan, Azzopardi, Leif, Aliannejadi, Mohammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Reusability of Conversational Search Test Collections
by: Abbasiantaeb, Zahra, et al.
Published: (2025)
by: Abbasiantaeb, Zahra, et al.
Published: (2025)
TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
TREC iKAT 2023: The Interactive Knowledge Assistance Track Overview
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
IRLab@iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets
by: Abbasiantaeb, Zahra, et al.
Published: (2025)
by: Abbasiantaeb, Zahra, et al.
Published: (2025)
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
by: Mousavian, Maryam, et al.
Published: (2025)
by: Mousavian, Maryam, et al.
Published: (2025)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Generating Multi-Aspect Queries for Conversational Search
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
Conversational Search: From Fundamentals to Frontiers in the LLM Era
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Re-Rankers as Relevance Judges
by: Meng, Chuan, et al.
Published: (2026)
by: Meng, Chuan, et al.
Published: (2026)
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
Total Recall QA: A Verifiable Evaluation Suite for Deep Research Agents
by: Rafiee, Mahta, et al.
Published: (2026)
by: Rafiee, Mahta, et al.
Published: (2026)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Improving Pinterest Search Relevance Using Large Language Models
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Information Farming: From Berry Picking to Berry Growing
by: Azzopardi, Leif, et al.
Published: (2026)
by: Azzopardi, Leif, et al.
Published: (2026)
Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrieval
by: Eichholtz, Arne, et al.
Published: (2026)
by: Eichholtz, Arne, et al.
Published: (2026)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
Investigating LLM Variability in Personalized Conversational Information Retrieval
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
Can Large Language Models Detect Rumors on Social Media?
by: Liu, Qiang, et al.
Published: (2024)
by: Liu, Qiang, et al.
Published: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
by: Yang, Jheng-Hong, et al.
Published: (2024)
by: Yang, Jheng-Hong, et al.
Published: (2024)
A Comparative Analysis of Faithfulness Metrics and Humans in Citation Evaluation
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Zero-Shot and Efficient Clarification Need Prediction in Conversational Search
by: Lu, Lili, et al.
Published: (2025)
by: Lu, Lili, et al.
Published: (2025)
Beyond Relevant Documents: A Knowledge-Intensive Approach for Query-Focused Summarization using Large Language Models
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)
by: de Jesus, Gabriel, et al.
Published: (2024)
Hallucination Detection and Evaluation of Large Language Model
by: Zhang, Chenggong, et al.
Published: (2025)
by: Zhang, Chenggong, et al.
Published: (2025)
Similar Items
-
Improving the Reusability of Conversational Search Test Collections
by: Abbasiantaeb, Zahra, et al.
Published: (2025) -
TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants
by: Aliannejadi, Mohammad, et al.
Published: (2024) -
TREC iKAT 2023: The Interactive Knowledge Assistance Track Overview
by: Aliannejadi, Mohammad, et al.
Published: (2024) -
IRLab@iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search
by: Lupart, Simon, et al.
Published: (2024) -
Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets
by: Abbasiantaeb, Zahra, et al.
Published: (2025)