LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
Fuente:
arXiv
Guardado en:
| Autores principales: | Takehi, Rikiya, Voorhees, Ellen M., Sakai, Tetsuya, Soboroff, Ian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Diversification as Risk Minimization
por: Takehi, Rikiya, et al.
Publicado: (2025)
por: Takehi, Rikiya, et al.
Publicado: (2025)
From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents
por: Akarsu, Meftun, et al.
Publicado: (2026)
por: Akarsu, Meftun, et al.
Publicado: (2026)
The Treatment of Ties in Rank-Biased Overlap
por: Corsi, Matteo, et al.
Publicado: (2024)
por: Corsi, Matteo, et al.
Publicado: (2024)
flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
por: Delmas, Damian
Publicado: (2026)
por: Delmas, Damian
Publicado: (2026)
Stop Using the Wilcoxon Test: Myth, Misconception and Misuse in IR Research
por: Urbano, Julián
Publicado: (2026)
por: Urbano, Julián
Publicado: (2026)
AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
por: Deng, Yangshen, et al.
Publicado: (2025)
por: Deng, Yangshen, et al.
Publicado: (2025)
Behavior-Aware Dual-Channel Preference Learning for Heterogeneous Sequential Recommendation
por: Xiao, Jing, et al.
Publicado: (2026)
por: Xiao, Jing, et al.
Publicado: (2026)
PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching
por: Cortès-Sebastià, Guillem, et al.
Publicado: (2025)
por: Cortès-Sebastià, Guillem, et al.
Publicado: (2025)
Siren Federate: Bridging document, relational, and graph models for exploratory graph analysis
por: Bordea, Georgeta, et al.
Publicado: (2025)
por: Bordea, Georgeta, et al.
Publicado: (2025)
Passing the Baton: High Throughput Distributed Disk-Based Vector Search with BatANN
por: Dang, Nam Anh, et al.
Publicado: (2025)
por: Dang, Nam Anh, et al.
Publicado: (2025)
A Recommender System Based on Binary Matrix Representations for Cognitive Disorders
por: Kutil, Raoul H., et al.
Publicado: (2025)
por: Kutil, Raoul H., et al.
Publicado: (2025)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
por: Schnabel, Julian A., et al.
Publicado: (2025)
por: Schnabel, Julian A., et al.
Publicado: (2025)
MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2026)
por: Türkmen, Mehmet Deniz, et al.
Publicado: (2026)
Relevance Filtering for Embedding-based Retrieval
por: Rossi, Nicholas, et al.
Publicado: (2024)
por: Rossi, Nicholas, et al.
Publicado: (2024)
Enhancing Relevance of Embedding-based Retrieval at Walmart
por: Lin, Juexin, et al.
Publicado: (2024)
por: Lin, Juexin, et al.
Publicado: (2024)
EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
por: Sun, Yuhong, et al.
Publicado: (2026)
por: Sun, Yuhong, et al.
Publicado: (2026)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
por: Keller, Jüri, et al.
Publicado: (2026)
por: Keller, Jüri, et al.
Publicado: (2026)
Leveraging LLMs to Enable Natural Language Search on Go-to-market Platforms
por: Yao, Jesse, et al.
Publicado: (2024)
por: Yao, Jesse, et al.
Publicado: (2024)
Diagnosing LLM-based Rerankers in Cold-Start Recommender Systems: Coverage, Exposure and Practical Mitigations
por: Lemdiasova, Ekaterina, et al.
Publicado: (2026)
por: Lemdiasova, Ekaterina, et al.
Publicado: (2026)
Where Relevance Emerges: A Layer-Wise Study of Internal Attention for Zero-Shot Re-Ranking
por: Chen, Haodong, et al.
Publicado: (2026)
por: Chen, Haodong, et al.
Publicado: (2026)
Beyond Similarity Search: A Unified Data Layer for Production RAG Systems
por: Budigi, Venkata Krishna Prasanth, et al.
Publicado: (2026)
por: Budigi, Venkata Krishna Prasanth, et al.
Publicado: (2026)
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
por: Dietz, Laura, et al.
Publicado: (2025)
por: Dietz, Laura, et al.
Publicado: (2025)
Generating Query Recommendations via LLMs
por: Bacciu, Andrea, et al.
Publicado: (2024)
por: Bacciu, Andrea, et al.
Publicado: (2024)
Boosting LLM-based Relevance Modeling with Distribution-Aware Robust Learning
por: Liu, Hong, et al.
Publicado: (2024)
por: Liu, Hong, et al.
Publicado: (2024)
Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality
por: Feng, Naihe, et al.
Publicado: (2025)
por: Feng, Naihe, et al.
Publicado: (2025)
LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
por: Abe, Kenya, et al.
Publicado: (2025)
por: Abe, Kenya, et al.
Publicado: (2025)
PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation
por: Ma, Chenglong, et al.
Publicado: (2025)
por: Ma, Chenglong, et al.
Publicado: (2025)
AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation
por: Yao, Zhihui, et al.
Publicado: (2026)
por: Yao, Zhihui, et al.
Publicado: (2026)
Criteria-Based LLM Relevance Judgments
por: Farzi, Naghmeh, et al.
Publicado: (2025)
por: Farzi, Naghmeh, et al.
Publicado: (2025)
Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
por: Ruggeri, Giuseppe, et al.
Publicado: (2025)
por: Ruggeri, Giuseppe, et al.
Publicado: (2025)
Memory Based Collaborative Filtering with Lucene
por: Gennaro, Claudio
Publicado: (2016)
por: Gennaro, Claudio
Publicado: (2016)
Unconstrained Monotonic Calibration of Predictions in Deep Ranking Systems
por: Bai, Yimeng, et al.
Publicado: (2025)
por: Bai, Yimeng, et al.
Publicado: (2025)
Context-Aware Lifelong Sequential Modeling for Online Click-Through Rate Prediction
por: Guo, Ting, et al.
Publicado: (2025)
por: Guo, Ting, et al.
Publicado: (2025)
LabelCraft: Empowering Short Video Recommendations with Automated Label Crafting
por: Bai, Yimeng, et al.
Publicado: (2023)
por: Bai, Yimeng, et al.
Publicado: (2023)
GradCraft: Elevating Multi-task Recommendations through Holistic Gradient Crafting
por: Bai, Yimeng, et al.
Publicado: (2024)
por: Bai, Yimeng, et al.
Publicado: (2024)
Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
por: Farzi, Naghmeh, et al.
Publicado: (2024)
por: Farzi, Naghmeh, et al.
Publicado: (2024)
Online Item Cold-Start Recommendation with Popularity-Aware Meta-Learning
por: Luo, Yunze, et al.
Publicado: (2024)
por: Luo, Yunze, et al.
Publicado: (2024)
Exploring Content-Based and Meta-Data Analysis for Detecting Fake News Infodemic: A case study on COVID-19
por: Ajao, Oluwaseun, et al.
Publicado: (2024)
por: Ajao, Oluwaseun, et al.
Publicado: (2024)
An Intrinsic Framework of Information Retrieval Evaluation Measures
por: Giner, Fernando
Publicado: (2023)
por: Giner, Fernando
Publicado: (2023)
Content-Aware Tweet Location Inference using Quadtree Spatial Partitioning and Jaccard-Cosine Word Embedding
por: Ajao, Oluwaseun, et al.
Publicado: (2024)
por: Ajao, Oluwaseun, et al.
Publicado: (2024)
Ejemplares similares
-
Diversification as Risk Minimization
por: Takehi, Rikiya, et al.
Publicado: (2025) -
From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents
por: Akarsu, Meftun, et al.
Publicado: (2026) -
The Treatment of Ties in Rank-Biased Overlap
por: Corsi, Matteo, et al.
Publicado: (2024) -
flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
por: Delmas, Damian
Publicado: (2026) -
Stop Using the Wilcoxon Test: Myth, Misconception and Misuse in IR Research
por: Urbano, Julián
Publicado: (2026)