LLM-based relevance assessment still can't replace human relevance assessment
Fuente:
arXiv
Saved in:
| Main Authors: | Clarke, Charles L. A., Dietz, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
by: Dietz, Laura, et al.
Published: (2025)
by: Dietz, Laura, et al.
Published: (2025)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
by: Lemdiasova, Ekaterina, et al.
Published: (2026)
Supporting Humans in Evaluating AI Summaries of Legal Depositions
by: Farzi, Naghmeh, et al.
Published: (2026)
by: Farzi, Naghmeh, et al.
Published: (2026)
LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
by: Abe, Kenya, et al.
Published: (2025)
by: Abe, Kenya, et al.
Published: (2025)
Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
by: Dietz, Laura, et al.
Published: (2026)
by: Dietz, Laura, et al.
Published: (2026)
MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval
by: Türkmen, Mehmet Deniz, et al.
Published: (2026)
by: Türkmen, Mehmet Deniz, et al.
Published: (2026)
Incorporating Q&A Nuggets into Retrieval-Augmented Generation
by: Dietz, Laura, et al.
Published: (2026)
by: Dietz, Laura, et al.
Published: (2026)
PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation
by: Ma, Chenglong, et al.
Published: (2025)
by: Ma, Chenglong, et al.
Published: (2025)
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
by: Takehi, Rikiya, et al.
Published: (2024)
by: Takehi, Rikiya, et al.
Published: (2024)
Cluster-based Graph Collaborative Filtering
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Relevance Filtering for Embedding-based Retrieval
by: Rossi, Nicholas, et al.
Published: (2024)
by: Rossi, Nicholas, et al.
Published: (2024)
A PLMs based protein retrieval framework
by: Wu, Yuxuan, et al.
Published: (2024)
by: Wu, Yuxuan, et al.
Published: (2024)
Enhancing Relevance of Embedding-based Retrieval at Walmart
by: Lin, Juexin, et al.
Published: (2024)
by: Lin, Juexin, et al.
Published: (2024)
Boosting LLM-based Relevance Modeling with Distribution-Aware Robust Learning
by: Liu, Hong, et al.
Published: (2024)
by: Liu, Hong, et al.
Published: (2024)
Can we predict QPP? An approach based on multivariate outliers
by: Chifu, Adrian-Gabriel, et al.
Published: (2024)
by: Chifu, Adrian-Gabriel, et al.
Published: (2024)
Explainable Graph-based Search for Lessons-Learned Documents in the Semiconductor Industry
by: Abu-Rasheed, Hasan, et al.
Published: (2021)
by: Abu-Rasheed, Hasan, et al.
Published: (2021)
ColBERT's [MASK]-based Query Augmentation: Effects of Quadrupling the Query Input Length
by: Giacalone, Ben, et al.
Published: (2024)
by: Giacalone, Ben, et al.
Published: (2024)
Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
by: Kuo, Tzu-Lin, et al.
Published: (2025)
by: Kuo, Tzu-Lin, et al.
Published: (2025)
Comparing the Utility, Preference, and Performance of Course Material Search Functionality and Retrieval-Augmented Generation Large Language Model (RAG-LLM) AI Chatbots in Information-Seeking Tasks
by: Pasquarelli, Leonardo, et al.
Published: (2024)
by: Pasquarelli, Leonardo, et al.
Published: (2024)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
by: Schnabel, Julian A., et al.
Published: (2025)
by: Schnabel, Julian A., et al.
Published: (2025)
Personalised Travel Recommendation based on Location Co-occurrence
by: Clements, Maarten, et al.
Published: (2011)
by: Clements, Maarten, et al.
Published: (2011)
Online Item Cold-Start Recommendation with Popularity-Aware Meta-Learning
by: Luo, Yunze, et al.
Published: (2024)
by: Luo, Yunze, et al.
Published: (2024)
Exploring Content-Based and Meta-Data Analysis for Detecting Fake News Infodemic: A case study on COVID-19
by: Ajao, Oluwaseun, et al.
Published: (2024)
by: Ajao, Oluwaseun, et al.
Published: (2024)
Content-Aware Tweet Location Inference using Quadtree Spatial Partitioning and Jaccard-Cosine Word Embedding
by: Ajao, Oluwaseun, et al.
Published: (2024)
by: Ajao, Oluwaseun, et al.
Published: (2024)
Text-like Encoding of Collaborative Information in Large Language Models for Recommendation
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Graph Reasoning for Explainable Cold Start Recommendation
by: Frej, Jibril, et al.
Published: (2024)
by: Frej, Jibril, et al.
Published: (2024)
MIND Your Language: A Multilingual Dataset for Cross-lingual News Recommendation
by: Iana, Andreea, et al.
Published: (2024)
by: Iana, Andreea, et al.
Published: (2024)
KamerRaad: Enhancing Information Retrieval in Belgian National Politics through Hierarchical Summarization and Conversational Interfaces
by: Rogiers, Alexander, et al.
Published: (2024)
by: Rogiers, Alexander, et al.
Published: (2024)
Generating Query Recommendations via LLMs
by: Bacciu, Andrea, et al.
Published: (2024)
by: Bacciu, Andrea, et al.
Published: (2024)
WindTunnel -- A Framework for Community Aware Sampling of Large Corpora
by: Iannelli, Michael
Published: (2024)
by: Iannelli, Michael
Published: (2024)
Future Impact Decomposition in Request-level Recommendations
by: Wang, Xiaobei, et al.
Published: (2024)
by: Wang, Xiaobei, et al.
Published: (2024)
Doc2Token: Bridging Vocabulary Gap by Predicting Missing Tokens for E-commerce Search
by: Li, Kaihao, et al.
Published: (2024)
by: Li, Kaihao, et al.
Published: (2024)
An Intrinsic Framework of Information Retrieval Evaluation Measures
by: Giner, Fernando
Published: (2023)
by: Giner, Fernando
Published: (2023)
Improving E-commerce Search with Category-Aligned Retrieval
by: Aliev, Rauf
Published: (2025)
by: Aliev, Rauf
Published: (2025)
Hyena Operator for Fast Sequential Recommendation
by: Liu, Jiahao, et al.
Published: (2026)
by: Liu, Jiahao, et al.
Published: (2026)
CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry
by: Ye, Xiaopeng, et al.
Published: (2025)
by: Ye, Xiaopeng, et al.
Published: (2025)
Information Retrieval for Climate Impact
by: de Rijke, Maarten, et al.
Published: (2025)
by: de Rijke, Maarten, et al.
Published: (2025)
Similar Items
-
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
by: Dietz, Laura, et al.
Published: (2025) -
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025) -
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
by: Farzi, Naghmeh, et al.
Published: (2024) -
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025) -
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
by: Lemdiasova, Ekaterina, et al.
Published: (2026)