Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gienapp, Lukas, Potthast, Martin, Yates, Andrew, Scells, Harrisen, Yang, Eugene |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
par: Gienapp, Lukas, et autres
Publié: (2024)
par: Gienapp, Lukas, et autres
Publié: (2024)
The Viability of Crowdsourcing for RAG Evaluation
par: Gienapp, Lukas, et autres
Publié: (2025)
par: Gienapp, Lukas, et autres
Publié: (2025)
AiReview: An Open Platform for Accelerating Systematic Reviews with LLMs
par: Mao, Xinyu, et autres
Publié: (2025)
par: Mao, Xinyu, et autres
Publié: (2025)
Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions
par: Heineking, Sebastian, et autres
Publié: (2024)
par: Heineking, Sebastian, et autres
Publié: (2024)
A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
par: Achkar, Pierre, et autres
Publié: (2026)
par: Achkar, Pierre, et autres
Publié: (2026)
Variations in Relevance Judgments and the Shelf Life of Test Collections
par: Parry, Andrew, et autres
Publié: (2025)
par: Parry, Andrew, et autres
Publié: (2025)
Zero-shot Generative Large Language Models for Systematic Review Screening Automation
par: Wang, Shuai, et autres
Publié: (2024)
par: Wang, Shuai, et autres
Publié: (2024)
Evaluating Generative Ad Hoc Information Retrieval
par: Gienapp, Lukas, et autres
Publié: (2023)
par: Gienapp, Lukas, et autres
Publié: (2023)
Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
par: Polyakov, Gregory, et autres
Publié: (2026)
par: Polyakov, Gregory, et autres
Publié: (2026)
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
par: Schlatt, Ferdinand, et autres
Publié: (2024)
par: Schlatt, Ferdinand, et autres
Publié: (2024)
DenseReviewer: A Screening Prioritisation Tool for Systematic Review based on Dense Retrieval
par: Mao, Xinyu, et autres
Publié: (2025)
par: Mao, Xinyu, et autres
Publié: (2025)
AutoBool: An Reinforcement-Learning trained LLM for Effective Automated Boolean Query Generation for Systematic Reviews
par: Wang, Shuai, et autres
Publié: (2025)
par: Wang, Shuai, et autres
Publié: (2025)
Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders
par: Schlatt, Ferdinand, et autres
Publié: (2024)
par: Schlatt, Ferdinand, et autres
Publié: (2024)
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
par: Wang, Shuai, et autres
Publié: (2025)
par: Wang, Shuai, et autres
Publié: (2025)
Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
par: Parry, Andrew, et autres
Publié: (2024)
par: Parry, Andrew, et autres
Publié: (2024)
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
par: Rottach, Florian, et autres
Publié: (2025)
par: Rottach, Florian, et autres
Publié: (2025)
Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage
par: Samuel, Saron, et autres
Publié: (2026)
par: Samuel, Saron, et autres
Publié: (2026)
Counterfactual Query Rewriting to Use Historical Relevance Feedback
par: Keller, Jüri, et autres
Publié: (2025)
par: Keller, Jüri, et autres
Publié: (2025)
Augmenting Researchy Questions with Sub-question Judgments
par: Ju, Jia-Huei, et autres
Publié: (2025)
par: Ju, Jia-Huei, et autres
Publié: (2025)
Is BERTopic Better than PLSA for Extracting Key Topics in Aviation Safety Reports?
par: Nanyonga, Aziida, et autres
Publié: (2025)
par: Nanyonga, Aziida, et autres
Publié: (2025)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
par: Rahmani, Hossein A., et autres
Publié: (2025)
par: Rahmani, Hossein A., et autres
Publié: (2025)
HLTCOE at LiveRAG: GPT-Researcher using ColBERT retrieval
par: Duh, Kevin, et autres
Publié: (2025)
par: Duh, Kevin, et autres
Publié: (2025)
RoutIR: Fast Serving of Retrieval Pipelines for Retrieval-Augmented Generation
par: Yang, Eugene, et autres
Publié: (2026)
par: Yang, Eugene, et autres
Publié: (2026)
Search for Coverage: Learning Coverage-Aware Retrieval with Augmented Sub-Question Answerability
par: Ju, Jia-Huei, et autres
Publié: (2026)
par: Ju, Jia-Huei, et autres
Publié: (2026)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
par: Rahmani, Hossein A., et autres
Publié: (2024)
par: Rahmani, Hossein A., et autres
Publié: (2024)
LANCER: LLM Reranking for Nugget Coverage
par: Ju, Jia-Huei, et autres
Publié: (2026)
par: Ju, Jia-Huei, et autres
Publié: (2026)
Are We Really Achieving Better Beyond-Accuracy Performance in Next Basket Recommendation?
par: Li, Ming, et autres
Publié: (2024)
par: Li, Ming, et autres
Publié: (2024)
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
par: Yu, Chuting, et autres
Publié: (2026)
par: Yu, Chuting, et autres
Publié: (2026)
Towards an In-Depth Comprehension of Case Relevance for Better Legal Retrieval
par: Li, Haitao, et autres
Publié: (2024)
par: Li, Haitao, et autres
Publié: (2024)
LLMJudge: LLMs for Relevance Judgments
par: Rahmani, Hossein A., et autres
Publié: (2024)
par: Rahmani, Hossein A., et autres
Publié: (2024)
Enhancing Health Information Retrieval with RAG by Prioritizing Topical Relevance and Factual Accuracy
par: Uapadhyay, Rishabh, et autres
Publié: (2025)
par: Uapadhyay, Rishabh, et autres
Publié: (2025)
Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
par: Nguyen, Thong, et autres
Publié: (2025)
par: Nguyen, Thong, et autres
Publié: (2025)
A Collection of Systematic Reviews in Computer Science
par: Achkar, Pierre, et autres
Publié: (2026)
par: Achkar, Pierre, et autres
Publié: (2026)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
par: Tamber, Manveer Singh, et autres
Publié: (2025)
par: Tamber, Manveer Singh, et autres
Publié: (2025)
Don't Use LLMs to Make Relevance Judgments
par: Soboroff, Ian
Publié: (2024)
par: Soboroff, Ian
Publié: (2024)
Characterising Topic Familiarity and Query Specificity Using Eye-Tracking Data
par: He, Jiaman, et autres
Publié: (2025)
par: He, Jiaman, et autres
Publié: (2025)
Towards Boosting LLMs-driven Relevance Modeling with Progressive Retrieved Behavior-augmented Prompting
par: Chen, Zeyuan, et autres
Publié: (2024)
par: Chen, Zeyuan, et autres
Publié: (2024)
Improving Topic Relevance Model by Mix-structured Summarization and LLM-based Data Augmentation
par: Liu, Yizhu, et autres
Publié: (2024)
par: Liu, Yizhu, et autres
Publié: (2024)
Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models
par: Choi, Jaekeol
Publié: (2024)
par: Choi, Jaekeol
Publié: (2024)
Hybrid Pooling with LLMs via Relevance Context Learning
par: Otero, David, et autres
Publié: (2026)
par: Otero, David, et autres
Publié: (2026)
Documents similaires
-
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
par: Gienapp, Lukas, et autres
Publié: (2024) -
The Viability of Crowdsourcing for RAG Evaluation
par: Gienapp, Lukas, et autres
Publié: (2025) -
AiReview: An Open Platform for Accelerating Systematic Reviews with LLMs
par: Mao, Xinyu, et autres
Publié: (2025) -
Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions
par: Heineking, Sebastian, et autres
Publié: (2024) -
A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
par: Achkar, Pierre, et autres
Publié: (2026)