An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
Fuente:
arXiv
Saved in:
| Main Authors: | Farzi, Naghmeh, Dietz, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Supporting Humans in Evaluating AI Summaries of Legal Depositions
by: Farzi, Naghmeh, et al.
Published: (2026)
by: Farzi, Naghmeh, et al.
Published: (2026)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Don't Use LLMs to Make Relevance Judgments
by: Soboroff, Ian
Published: (2024)
by: Soboroff, Ian
Published: (2024)
Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage
by: Samuel, Saron, et al.
Published: (2026)
by: Samuel, Saron, et al.
Published: (2026)
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)
by: de Jesus, Gabriel, et al.
Published: (2024)
Variations in Relevance Judgments and the Shelf Life of Test Collections
by: Parry, Andrew, et al.
Published: (2025)
by: Parry, Andrew, et al.
Published: (2025)
A Workbench for Autograding Retrieve/Generate Systems
by: Dietz, Laura
Published: (2024)
by: Dietz, Laura
Published: (2024)
A Study of Relevance Judgments.
by: Cuadra, Carlos A.
Published: (1968)
by: Cuadra, Carlos A.
Published: (1968)
TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models
by: Zhou, Jianqun, et al.
Published: (2024)
by: Zhou, Jianqun, et al.
Published: (2024)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
by: Ma, Shengjie, et al.
Published: (2024)
by: Ma, Shengjie, et al.
Published: (2024)
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
by: Zhao, Xinran, et al.
Published: (2024)
by: Zhao, Xinran, et al.
Published: (2024)
LLM-Driven Usefulness Judgment for Web Search Evaluation
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
Can We Use Large Language Models to Fill Relevance Judgment Holes?
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
by: Abbasiantaeb, Zahra, et al.
Published: (2024)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
LLM-based relevance assessment still can't replace human relevance assessment
by: Clarke, Charles L. A., et al.
Published: (2024)
by: Clarke, Charles L. A., et al.
Published: (2024)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
by: Keller, Jüri, et al.
Published: (2026)
by: Keller, Jüri, et al.
Published: (2026)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Beyond Relevance: An Adaptive Exploration-Based Framework for Personalized Recommendations
by: Bianchi, Edoardo
Published: (2025)
by: Bianchi, Edoardo
Published: (2025)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
Beyond Relevance: Improving User Engagement by Personalization for Short-Video Search
by: Bao, Wentian, et al.
Published: (2024)
by: Bao, Wentian, et al.
Published: (2024)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
The Cranfield II Relevance Assessments: A Critical Evaluation
by: Harter, Stephen P.
Published: (1971)
by: Harter, Stephen P.
Published: (1971)
A Deep Learning Approach for Selective Relevance Feedback
by: Datta, Suchana, et al.
Published: (2024)
by: Datta, Suchana, et al.
Published: (2024)
Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative Approach
by: Qin, Weicong, et al.
Published: (2023)
by: Qin, Weicong, et al.
Published: (2023)
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
by: Christakopoulou, Evangelia, et al.
Published: (2026)
by: Christakopoulou, Evangelia, et al.
Published: (2026)
A Comprehensive Review on Hashtag Recommendation: From Traditional to Deep Learning and Beyond
by: Bansal, Shubhi, et al.
Published: (2025)
by: Bansal, Shubhi, et al.
Published: (2025)
Prism-Reranker: Beyond Relevance Scoring -- Jointly Producing Contributions and Evidence for Agentic Retrieval
by: Zhang, Dun
Published: (2026)
by: Zhang, Dun
Published: (2026)
Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models
by: Choi, Jaekeol
Published: (2024)
by: Choi, Jaekeol
Published: (2024)
Similar Items
-
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025) -
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024) -
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025) -
Supporting Humans in Evaluating AI Summaries of Legal Depositions
by: Farzi, Naghmeh, et al.
Published: (2026) -
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)