LLMs Can Patch Up Missing Relevance Judgments in Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Upadhyay, Shivani, Kamalloo, Ehsan, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2025)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2025)
LLMJudge: LLMs for Relevance Judgments
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2025)
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2025)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2024)
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2024)
Don't Use LLMs to Make Relevance Judgments
von: Soboroff, Ian
Veröffentlicht: (2024)
von: Soboroff, Ian
Veröffentlicht: (2024)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
von: Jedidi, Nour, et al.
Veröffentlicht: (2026)
von: Jedidi, Nour, et al.
Veröffentlicht: (2026)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
von: Thakur, Nandan, et al.
Veröffentlicht: (2023)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2026)
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2026)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
von: Farzi, Naghmeh, et al.
Veröffentlicht: (2024)
von: Farzi, Naghmeh, et al.
Veröffentlicht: (2024)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
von: Thakur, Nandan, et al.
Veröffentlicht: (2025)
Can We Use Large Language Models to Fill Relevance Judgment Holes?
von: Abbasiantaeb, Zahra, et al.
Veröffentlicht: (2024)
von: Abbasiantaeb, Zahra, et al.
Veröffentlicht: (2024)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
von: de Jesus, Gabriel, et al.
Veröffentlicht: (2024)
von: de Jesus, Gabriel, et al.
Veröffentlicht: (2024)
Variations in Relevance Judgments and the Shelf Life of Test Collections
von: Parry, Andrew, et al.
Veröffentlicht: (2025)
von: Parry, Andrew, et al.
Veröffentlicht: (2025)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Benchmarking LLM-based Relevance Judgment Methods
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
A Study of Relevance Judgments.
von: Cuadra, Carlos A.
Veröffentlicht: (1968)
von: Cuadra, Carlos A.
Veröffentlicht: (1968)
TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
von: Dewan, Mouly, et al.
Veröffentlicht: (2025)
von: Dewan, Mouly, et al.
Veröffentlicht: (2025)
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
von: Iturra-Bocaz, Gabriel, et al.
Veröffentlicht: (2025)
von: Iturra-Bocaz, Gabriel, et al.
Veröffentlicht: (2025)
Query-driven Relevant Paragraph Extraction from Legal Judgments
von: Santosh, T. Y. S. S, et al.
Veröffentlicht: (2024)
von: Santosh, T. Y. S. S, et al.
Veröffentlicht: (2024)
Can We Trust Recommender System Fairness Evaluation? The Role of Fairness and Relevance
von: Rampisela, Theresia Veronika, et al.
Veröffentlicht: (2024)
von: Rampisela, Theresia Veronika, et al.
Veröffentlicht: (2024)
Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2024)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
von: Zhang, Hengran, et al.
Veröffentlicht: (2024)
von: Zhang, Hengran, et al.
Veröffentlicht: (2024)
Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
von: Lin, Jimmy
Veröffentlicht: (2024)
von: Lin, Jimmy
Veröffentlicht: (2024)
LLM-Driven Usefulness Judgment for Web Search Evaluation
von: Dewan, Mouly, et al.
Veröffentlicht: (2025)
von: Dewan, Mouly, et al.
Veröffentlicht: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2025)
Criteria-Based LLM Relevance Judgments
von: Farzi, Naghmeh, et al.
Veröffentlicht: (2025)
von: Farzi, Naghmeh, et al.
Veröffentlicht: (2025)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
von: Keller, Jüri, et al.
Veröffentlicht: (2026)
von: Keller, Jüri, et al.
Veröffentlicht: (2026)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2025)
von: Fernandes, Leandro Carísio, et al.
Veröffentlicht: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
von: Upadhyay, Shivani, et al.
Veröffentlicht: (2025) -
LLMJudge: LLMs for Relevance Judgments
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024) -
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2025) -
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
von: Thakur, Nandan, et al.
Veröffentlicht: (2024) -
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
von: Sharifymoghaddam, Sahel, et al.
Veröffentlicht: (2024)