Don't Use LLMs to Make Relevance Judgments
Fuente:
arXiv
Saved in:
| Main Author: | Soboroff, Ian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
by: Takehi, Rikiya, et al.
Published: (2024)
by: Takehi, Rikiya, et al.
Published: (2024)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)
by: de Jesus, Gabriel, et al.
Published: (2024)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Variations in Relevance Judgments and the Shelf Life of Test Collections
by: Parry, Andrew, et al.
Published: (2025)
by: Parry, Andrew, et al.
Published: (2025)
Can We Use Large Language Models to Fill Relevance Judgment Holes?
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Rank, Don't Generate: Statement-level Ranking for Explainable Recommendation
by: Kabongo, Ben, et al.
Published: (2026)
by: Kabongo, Ben, et al.
Published: (2026)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
A Study of Relevance Judgments.
by: Cuadra, Carlos A.
Published: (1968)
by: Cuadra, Carlos A.
Published: (1968)
TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
Don't Start Over: A Cost-Effective Framework for Migrating Personalized Prompts Between LLMs
by: Zhao, Ziyi, et al.
Published: (2026)
by: Zhao, Ziyi, et al.
Published: (2026)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Don't Click the Bait: Title Debiasing News Recommendation via Cross-Field Contrastive Learning
by: Shu, Yijie, et al.
Published: (2024)
by: Shu, Yijie, et al.
Published: (2024)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
by: Ma, Shengjie, et al.
Published: (2024)
by: Ma, Shengjie, et al.
Published: (2024)
Don't Measure Once: Measuring Visibility in AI Search (GEO)
by: Schulte, Julius, et al.
Published: (2026)
by: Schulte, Julius, et al.
Published: (2026)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
by: Keller, Jüri, et al.
Published: (2026)
by: Keller, Jüri, et al.
Published: (2026)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Counterfactual Query Rewriting to Use Historical Relevance Feedback
by: Keller, Jüri, et al.
Published: (2025)
by: Keller, Jüri, et al.
Published: (2025)
Hybrid Pooling with LLMs via Relevance Context Learning
by: Otero, David, et al.
Published: (2026)
by: Otero, David, et al.
Published: (2026)
CoverageBench: Evaluating Information Coverage across Tasks and Domains
by: Samuel, Saron, et al.
Published: (2026)
by: Samuel, Saron, et al.
Published: (2026)
"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation
by: Di Scala, Daan, et al.
Published: (2026)
by: Di Scala, Daan, et al.
Published: (2026)
Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network
by: Chen, Zheyu, et al.
Published: (2024)
by: Chen, Zheyu, et al.
Published: (2024)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
Expanding Relevance Judgments for Medical Case-based Retrieval Task with Multimodal LLMs
by: Pires, Catarina, et al.
Published: (2025)
by: Pires, Catarina, et al.
Published: (2025)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
You Don't Bring Me Flowers: Mitigating Unwanted Recommendations Through Conformal Risk Control
by: De Toni, Giovanni, et al.
Published: (2025)
by: De Toni, Giovanni, et al.
Published: (2025)
Don't Waste It: Guiding Generative Recommenders with Structured Human Priors via Multi-Head Decoding
by: Zhang, Yunkai, et al.
Published: (2025)
by: Zhang, Yunkai, et al.
Published: (2025)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
Similar Items
-
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024) -
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
by: Takehi, Rikiya, et al.
Published: (2024) -
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024) -
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024) -
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)