Generative Information Retrieval Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Alaofi, Marwah, Arabzadeh, Negar, Clarke, Charles L. A., Sanderson, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
LLMs can be Fooled into Labelling a Document as Relevant (best café near me; this paper is perfectly relevant)
by: Alaofi, Marwah, et al.
Published: (2025)
by: Alaofi, Marwah, et al.
Published: (2025)
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
by: Alaofi, Marwah, et al.
Published: (2025)
by: Alaofi, Marwah, et al.
Published: (2025)
Demographically-Inspired Query Variants Using an LLM
by: Alaofi, Marwah, et al.
Published: (2025)
by: Alaofi, Marwah, et al.
Published: (2025)
Benchmarking LLM-based Relevance Judgment Methods
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
by: Arabzadeh, Negar, et al.
Published: (2025)
by: Arabzadeh, Negar, et al.
Published: (2025)
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
by: Bigdeli, Amin, et al.
Published: (2024)
by: Bigdeli, Amin, et al.
Published: (2024)
Personalisation of Generic Library Search Results Using Student Enrolment Information
by: Alaofi, Marwah, et al.
Published: (2015)
by: Alaofi, Marwah, et al.
Published: (2015)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Offline Evaluation of Set-Based Text-to-Image Generation
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
by: Arabzadeh, Ardalan
Published: (2025)
by: Arabzadeh, Ardalan
Published: (2025)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2025)
by: Bigdeli, Amin, et al.
Published: (2025)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
A Reproducibility Study of LLM-Based Query Reformulation
by: Bigdeli, Amin, et al.
Published: (2026)
by: Bigdeli, Amin, et al.
Published: (2026)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
by: Amirshahi, Shakiba, et al.
Published: (2025)
by: Amirshahi, Shakiba, et al.
Published: (2025)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
by: Ebrahimi, Sajad, et al.
Published: (2025)
by: Ebrahimi, Sajad, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Metamorphic Evaluation of ChatGPT as a Recommender System
by: Khirbat, Madhurima, et al.
Published: (2024)
by: Khirbat, Madhurima, et al.
Published: (2024)
Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Annotative Indexing
by: Clarke, Charles L. A.
Published: (2024)
by: Clarke, Charles L. A.
Published: (2024)
Evaluating and Addressing Fairness Across User Groups in Negative Sampling for Recommender Systems
by: Xuan, Yueqing, et al.
Published: (2023)
by: Xuan, Yueqing, et al.
Published: (2023)
LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
by: Dietz, Laura, et al.
Published: (2025)
by: Dietz, Laura, et al.
Published: (2025)
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
by: Zhang, Dake, et al.
Published: (2026)
by: Zhang, Dake, et al.
Published: (2026)
Online and Offline Evaluation in Search Clarification
by: Tavakoli, Leila, et al.
Published: (2024)
by: Tavakoli, Leila, et al.
Published: (2024)
WildClaims: Information Access Conversations in the Wild(Chat)
by: Joko, Hideaki, et al.
Published: (2025)
by: Joko, Hideaki, et al.
Published: (2025)
Peerispect: Claim Verification in Scientific Peer Reviews
by: Ghorbanpour, Ali, et al.
Published: (2026)
by: Ghorbanpour, Ali, et al.
Published: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
Diversity-Augmented Negative Sampling for Implicit Collaborative Filtering
by: Xuan, Yueqing, et al.
Published: (2025)
by: Xuan, Yueqing, et al.
Published: (2025)
Beyond Utility: Evaluating LLM as Recommender
by: Jiang, Chumeng, et al.
Published: (2024)
by: Jiang, Chumeng, et al.
Published: (2024)
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)
by: Gienapp, Lukas, et al.
Published: (2023)
Green Recommender Systems: Optimizing Dataset Size for Energy-Efficient Algorithm Performance
by: Arabzadeh, Ardalan, et al.
Published: (2024)
by: Arabzadeh, Ardalan, et al.
Published: (2024)
Replicability Measures for Longitudinal Information Retrieval Evaluation
by: Keller, Jüri, et al.
Published: (2024)
by: Keller, Jüri, et al.
Published: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Interactions with Generative Information Retrieval Systems
by: Aliannejadi, Mohammad, et al.
Published: (2024)
by: Aliannejadi, Mohammad, et al.
Published: (2024)
On the Robustness of Generative Information Retrieval Models
by: Liu, Yu-An, et al.
Published: (2024)
by: Liu, Yu-An, et al.
Published: (2024)
Retrieval Augmented Generation Evaluation for Health Documents
by: Ceresa, Mario, et al.
Published: (2025)
by: Ceresa, Mario, et al.
Published: (2025)
Evaluating D-MERIT of Partial-annotation on Information Retrieval
by: Rassin, Royi, et al.
Published: (2024)
by: Rassin, Royi, et al.
Published: (2024)
Similar Items
-
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
by: Arabzadeh, Negar, et al.
Published: (2024) -
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
by: Arabzadeh, Negar, et al.
Published: (2024) -
A Comparison of Methods for Evaluating Generative IR
by: Arabzadeh, Negar, et al.
Published: (2024) -
LLMs can be Fooled into Labelling a Document as Relevant (best café near me; this paper is perfectly relevant)
by: Alaofi, Marwah, et al.
Published: (2025) -
Can Generative LLMs Create Query Variants for Test Collections? An Exploratory Study
by: Alaofi, Marwah, et al.
Published: (2025)