TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Dewan, Mouly, Liu, Jiqun, Shah, Chirag |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Driven Usefulness Judgment for Web Search Evaluation
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
LLM-Driven Usefulness Labeling for IR Evaluation
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
Mind Over Misinformation: Investigating the Factors of Cognitive Influences in Information Acceptance
by: Mouly Dewan, et al.
Published: (2024)
by: Mouly Dewan, et al.
Published: (2024)
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Improving Data Reusability in Interactive Information Retrieval: Insights from the Community
by: Jiang, Tianji, et al.
Published: (2025)
by: Jiang, Tianji, et al.
Published: (2025)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability
by: Jiang, Tianji, et al.
Published: (2024)
by: Jiang, Tianji, et al.
Published: (2024)
LLMJudge: LLMs for Relevance Judgments
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
by: Chen, Nuo, et al.
Published: (2026)
by: Chen, Nuo, et al.
Published: (2026)
Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval
by: Ma, Shengjie, et al.
Published: (2024)
by: Ma, Shengjie, et al.
Published: (2024)
Don't Use LLMs to Make Relevance Judgments
by: Soboroff, Ian
Published: (2024)
by: Soboroff, Ian
Published: (2024)
Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Variations in Relevance Judgments and the Shelf Life of Test Collections
by: Parry, Andrew, et al.
Published: (2025)
by: Parry, Andrew, et al.
Published: (2025)
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Exploring Large Language Models for Relevance Judgments in Tetun
by: de Jesus, Gabriel, et al.
Published: (2024)
by: de Jesus, Gabriel, et al.
Published: (2024)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
An Exam-based Evaluation Approach Beyond Traditional Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2024)
by: Farzi, Naghmeh, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Beyond Static Evaluation: Rethinking the Assessment of Personalized Agent Adaptability in Information Retrieval
by: Kaur, Kirandeep, et al.
Published: (2025)
by: Kaur, Kirandeep, et al.
Published: (2025)
Reproducing NevIR: Negation in Neural Information Retrieval
by: Elsen, Coen van den, et al.
Published: (2025)
by: Elsen, Coen van den, et al.
Published: (2025)
A Study of Relevance Judgments.
by: Cuadra, Carlos A.
Published: (1968)
by: Cuadra, Carlos A.
Published: (1968)
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
by: Iturra-Bocaz, Gabriel, et al.
Published: (2025)
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
by: Keller, Jüri, et al.
Published: (2026)
by: Keller, Jüri, et al.
Published: (2026)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
Reproducing Complex Set-Compositional Information Retrieval
by: Degenhart, Vincent, et al.
Published: (2026)
by: Degenhart, Vincent, et al.
Published: (2026)
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
by: Christakopoulou, Evangelia, et al.
Published: (2026)
by: Christakopoulou, Evangelia, et al.
Published: (2026)
A Reproducibility Study of Metacognitive Retrieval-Augmented Generation
by: Iturra-Bocaz, Gabriel, et al.
Published: (2026)
by: Iturra-Bocaz, Gabriel, et al.
Published: (2026)
Query-driven Relevant Paragraph Extraction from Legal Judgments
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
A Case-Driven Multi-Agent Framework for E-Commerce Search Relevance
by: Commerce Search Relevance Team
Published: (2026)
by: Commerce Search Relevance Team
Published: (2026)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Cognitively Biased Users Interacting with Algorithmically Biased Results in Whole-Session Search on Debated Topics
by: Wang, Ben, et al.
Published: (2024)
by: Wang, Ben, et al.
Published: (2024)
Enhancing Health Information Retrieval with RAG by Prioritizing Topical Relevance and Factual Accuracy
by: Uapadhyay, Rishabh, et al.
Published: (2025)
by: Uapadhyay, Rishabh, et al.
Published: (2025)
Discovering Biases in Information Retrieval Models Using Relevance Thesaurus as Global Explanation
by: Kim, Youngwoo, et al.
Published: (2024)
by: Kim, Youngwoo, et al.
Published: (2024)
A Reproducibility Study of Graph-Based Legal Case Retrieval
by: Donabauer, Gregor, et al.
Published: (2025)
by: Donabauer, Gregor, et al.
Published: (2025)
Panmodal Information Interaction
by: Shah, Chirag, et al.
Published: (2024)
by: Shah, Chirag, et al.
Published: (2024)
Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
On the Reproducibility of Learned Sparse Retrieval Adaptations for Long Documents
by: Lionis, Emmanouil Georgios, et al.
Published: (2025)
by: Lionis, Emmanouil Georgios, et al.
Published: (2025)
Similar Items
-
LLM-Driven Usefulness Judgment for Web Search Evaluation
by: Dewan, Mouly, et al.
Published: (2025) -
LLM-Driven Usefulness Labeling for IR Evaluation
by: Dewan, Mouly, et al.
Published: (2025) -
Mind Over Misinformation: Investigating the Factors of Cognitive Influences in Information Acceptance
by: Mouly Dewan, et al.
Published: (2024) -
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
by: Chen, Nuo, et al.
Published: (2025) -
Improving Data Reusability in Interactive Information Retrieval: Insights from the Community
by: Jiang, Tianji, et al.
Published: (2025)