Best in Tau@LLMJudge: Criteria-Based Relevance Evaluation with Llama3
Fuente:
arXiv
Guardado en:
| Autores principales: | Farzi, Naghmeh, Dietz, Laura |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Criteria-Based LLM Relevance Judgments
por: Farzi, Naghmeh, et al.
Publicado: (2025)
por: Farzi, Naghmeh, et al.
Publicado: (2025)
Does UMBRELA Work on Other LLMs?
por: Farzi, Naghmeh, et al.
Publicado: (2025)
por: Farzi, Naghmeh, et al.
Publicado: (2025)
Supporting Humans in Evaluating AI Summaries of Legal Depositions
por: Farzi, Naghmeh, et al.
Publicado: (2026)
por: Farzi, Naghmeh, et al.
Publicado: (2026)
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
por: Dietz, Laura, et al.
Publicado: (2025)
por: Dietz, Laura, et al.
Publicado: (2025)
Thinking Broad, Acting Fast: Latent Reasoning Distillation from Multi-Perspective Chain-of-Thought for E-Commerce Relevance
por: Qiu, Baopu, et al.
Publicado: (2026)
por: Qiu, Baopu, et al.
Publicado: (2026)
AI Enhanced Ontology Driven NLP for Intelligent Cloud Resource Query Processing Using Knowledge Graphs
por: Sunkara, Krishna Chaitanya, et al.
Publicado: (2025)
por: Sunkara, Krishna Chaitanya, et al.
Publicado: (2025)
Action-Aware Generative Sequence Modeling for Short Video Recommendation
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
Effective Diversification of Multi-Carousel Book Recommendation
por: Wilten, Daniël, et al.
Publicado: (2025)
por: Wilten, Daniël, et al.
Publicado: (2025)
MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
por: Zhang, Ranxu, et al.
Publicado: (2026)
por: Zhang, Ranxu, et al.
Publicado: (2026)
Large Language Models for Relevance Judgment in Product Search
por: Mehrdad, Navid, et al.
Publicado: (2024)
por: Mehrdad, Navid, et al.
Publicado: (2024)
Expanding Relevance Judgments for Medical Case-based Retrieval Task with Multimodal LLMs
por: Pires, Catarina, et al.
Publicado: (2025)
por: Pires, Catarina, et al.
Publicado: (2025)
Relevance Filtering for Embedding-based Retrieval
por: Rossi, Nicholas, et al.
Publicado: (2024)
por: Rossi, Nicholas, et al.
Publicado: (2024)
Re-Rankers as Relevance Judges
por: Meng, Chuan, et al.
Publicado: (2026)
por: Meng, Chuan, et al.
Publicado: (2026)
Enhancing Relevance of Embedding-based Retrieval at Walmart
por: Lin, Juexin, et al.
Publicado: (2024)
por: Lin, Juexin, et al.
Publicado: (2024)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
por: Keller, Jüri, et al.
Publicado: (2026)
por: Keller, Jüri, et al.
Publicado: (2026)
MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models
por: Li, Haoxuan, et al.
Publicado: (2024)
por: Li, Haoxuan, et al.
Publicado: (2024)
Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval
por: Mekonnen, Kidist Amde, et al.
Publicado: (2025)
por: Mekonnen, Kidist Amde, et al.
Publicado: (2025)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
por: Farzi, Naghmeh, et al.
Publicado: (2024)
por: Farzi, Naghmeh, et al.
Publicado: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
por: Meng, Chuan, et al.
Publicado: (2024)
por: Meng, Chuan, et al.
Publicado: (2024)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
por: Schnabel, Julian A., et al.
Publicado: (2025)
por: Schnabel, Julian A., et al.
Publicado: (2025)
Where Relevance Emerges: A Layer-Wise Study of Internal Attention for Zero-Shot Re-Ranking
por: Chen, Haodong, et al.
Publicado: (2026)
por: Chen, Haodong, et al.
Publicado: (2026)
flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
por: Delmas, Damian
Publicado: (2026)
por: Delmas, Damian
Publicado: (2026)
Modeling Stage-wise Evolution of User Interests for News Recommendation
por: Cheng, Zhiyong, et al.
Publicado: (2026)
por: Cheng, Zhiyong, et al.
Publicado: (2026)
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
por: Takehi, Rikiya, et al.
Publicado: (2024)
por: Takehi, Rikiya, et al.
Publicado: (2024)
Counterfactual Evaluation of Ads Ranking Models through Domain Adaptation
por: Radwan, Mohamed A., et al.
Publicado: (2024)
por: Radwan, Mohamed A., et al.
Publicado: (2024)
Evaluating the Effectiveness of Large Language Models in Automated News Article Summarization
por: Houamegni, Lionel Richy Panlap, et al.
Publicado: (2025)
por: Houamegni, Lionel Richy Panlap, et al.
Publicado: (2025)
Peeling Back the Layers: An In-Depth Evaluation of Encoder Architectures in Neural News Recommenders
por: Iana, Andreea, et al.
Publicado: (2024)
por: Iana, Andreea, et al.
Publicado: (2024)
Exploring Information Retrieval Landscapes: An Investigation of a Novel Evaluation Techniques and Comparative Document Splitting Methods
por: Narimissa, Esmaeil, et al.
Publicado: (2024)
por: Narimissa, Esmaeil, et al.
Publicado: (2024)
OmniSearchSage: Multi-Task Multi-Entity Embeddings for Pinterest Search
por: Agarwal, Prabhat, et al.
Publicado: (2024)
por: Agarwal, Prabhat, et al.
Publicado: (2024)
PyTerrier-GenRank: The PyTerrier Plugin for Reranking with Large Language Models
por: Dhole, Kaustubh D.
Publicado: (2024)
por: Dhole, Kaustubh D.
Publicado: (2024)
LiRank: Industrial Large Scale Ranking Models at LinkedIn
por: Borisyuk, Fedor, et al.
Publicado: (2024)
por: Borisyuk, Fedor, et al.
Publicado: (2024)
Pfeed: Generating near real-time personalized feeds using precomputed embedding similarities
por: Gebre, Binyam, et al.
Publicado: (2024)
por: Gebre, Binyam, et al.
Publicado: (2024)
A Comparative Study of Hybrid Models in Health Misinformation Text Classification
por: Sikosana, Mkululi, et al.
Publicado: (2024)
por: Sikosana, Mkululi, et al.
Publicado: (2024)
Improving BM25 Code Retrieval Under Fixed Generic Tokenization: Adaptive q-Log Odds as a Drop-In BM25 Fix
por: Radha, Santosh Kumar, et al.
Publicado: (2026)
por: Radha, Santosh Kumar, et al.
Publicado: (2026)
Reinforcing User Interest Evolution in Multi-Scenario Learning for recommender systems
por: Feng, Zhijian, et al.
Publicado: (2025)
por: Feng, Zhijian, et al.
Publicado: (2025)
InteractRank: Personalized Web-Scale Search Pre-Ranking with Cross Interaction Features
por: Khandagale, Sujay, et al.
Publicado: (2025)
por: Khandagale, Sujay, et al.
Publicado: (2025)
Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language Models
por: Ahn, Seokho, et al.
Publicado: (2026)
por: Ahn, Seokho, et al.
Publicado: (2026)
Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering
por: Chen, Haodong, et al.
Publicado: (2025)
por: Chen, Haodong, et al.
Publicado: (2025)
Scientific and Technological Information Oriented Semantics-adversarial and Media-adversarial Cross-media Retrieval
por: Li, Ang, et al.
Publicado: (2022)
por: Li, Ang, et al.
Publicado: (2022)
ReCast: Recasting Learning Signals for Reinforcement Learning in Generative Recommendation
por: Zhang, Peiyan, et al.
Publicado: (2026)
por: Zhang, Peiyan, et al.
Publicado: (2026)
Ejemplares similares
-
Criteria-Based LLM Relevance Judgments
por: Farzi, Naghmeh, et al.
Publicado: (2025) -
Does UMBRELA Work on Other LLMs?
por: Farzi, Naghmeh, et al.
Publicado: (2025) -
Supporting Humans in Evaluating AI Summaries of Legal Depositions
por: Farzi, Naghmeh, et al.
Publicado: (2026) -
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
por: Dietz, Laura, et al.
Publicado: (2025) -
Thinking Broad, Acting Fast: Latent Reasoning Distillation from Multi-Perspective Chain-of-Thought for E-Commerce Relevance
por: Qiu, Baopu, et al.
Publicado: (2026)