A Comparison of Methods for Evaluating Generative IR
Fuente:
arXiv
Saved in:
| Main Authors: | Arabzadeh, Negar, Clarke, Charles L. A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024)
by: Alaofi, Marwah, et al.
Published: (2024)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2024)
by: Bigdeli, Amin, et al.
Published: (2024)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
by: Arabzadeh, Ardalan
Published: (2025)
by: Arabzadeh, Ardalan
Published: (2025)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
by: Ebrahimi, Sajad, et al.
Published: (2025)
by: Ebrahimi, Sajad, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Annotative Indexing
by: Clarke, Charles L. A.
Published: (2024)
by: Clarke, Charles L. A.
Published: (2024)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
by: Amirshahi, Shakiba, et al.
Published: (2025)
by: Amirshahi, Shakiba, et al.
Published: (2025)
Peerispect: Claim Verification in Scientific Peer Reviews
by: Ghorbanpour, Ali, et al.
Published: (2026)
by: Ghorbanpour, Ali, et al.
Published: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
Beyond Utility: Evaluating LLM as Recommender
by: Jiang, Chumeng, et al.
Published: (2024)
by: Jiang, Chumeng, et al.
Published: (2024)
Green Recommender Systems: Optimizing Dataset Size for Energy-Efficient Algorithm Performance
by: Arabzadeh, Ardalan, et al.
Published: (2024)
by: Arabzadeh, Ardalan, et al.
Published: (2024)
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
by: Zhang, Dake, et al.
Published: (2026)
by: Zhang, Dake, et al.
Published: (2026)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
LLM-based relevance assessment still can't replace human relevance assessment
by: Clarke, Charles L. A., et al.
Published: (2024)
by: Clarke, Charles L. A., et al.
Published: (2024)
ASPIRE: Assistive System for Performance Evaluation in IR
by: Peikos, Georgios, et al.
Published: (2024)
by: Peikos, Georgios, et al.
Published: (2024)
Evaluation of Temporal Change in IR Test Collections
by: Keller, Jüri, et al.
Published: (2024)
by: Keller, Jüri, et al.
Published: (2024)
LLM-Driven Usefulness Labeling for IR Evaluation
by: Dewan, Mouly, et al.
Published: (2025)
by: Dewan, Mouly, et al.
Published: (2025)
WildClaims: Information Access Conversations in the Wild(Chat)
by: Joko, Hideaki, et al.
Published: (2025)
by: Joko, Hideaki, et al.
Published: (2025)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
ir_explain: a Python Library of Explainable IR Methods
by: Saha, Sourav, et al.
Published: (2024)
by: Saha, Sourav, et al.
Published: (2024)
Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
by: Cancellieri, Matteo, et al.
Published: (2025)
by: Cancellieri, Matteo, et al.
Published: (2025)
Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR
by: Datta, Suchana, et al.
Published: (2026)
by: Datta, Suchana, et al.
Published: (2026)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
by: Vast, Mathias, et al.
Published: (2024)
by: Vast, Mathias, et al.
Published: (2024)
RoutIR: Fast Serving of Retrieval Pipelines for Retrieval-Augmented Generation
by: Yang, Eugene, et al.
Published: (2026)
by: Yang, Eugene, et al.
Published: (2026)
SimEval-IR: A Unified Toolkit and Benchmark Suite for Evaluating User Simulators and Search Sessions
by: Zerhoudi, Saber
Published: (2026)
by: Zerhoudi, Saber
Published: (2026)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Systems on Web and Scientific Data
by: Cancellieri, Matteo, et al.
Published: (2025)
by: Cancellieri, Matteo, et al.
Published: (2025)
Dissertation: On the Theoretical Foundation of Model Comparison and Evaluation for Recommender System
by: Li, Dong
Published: (2024)
by: Li, Dong
Published: (2024)
Genetic Approach to Mitigate Hallucination in Generative IR
by: Kulkarni, Hrishikesh, et al.
Published: (2024)
by: Kulkarni, Hrishikesh, et al.
Published: (2024)
Similar Items
-
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024) -
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024) -
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025) -
Generative Information Retrieval Evaluation
by: Alaofi, Marwah, et al.
Published: (2024) -
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)