Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arabzadeh, Negar, Bigdeli, Amin, Clarke, Charles L. A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
von: Bigdeli, Amin, et al.
Veröffentlicht: (2024)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2024)
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Adversarial Attacks against Neural Ranking Models via In-Context Learning
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
A Comparison of Methods for Evaluating Generative IR
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Generative Information Retrieval Evaluation
von: Alaofi, Marwah, et al.
Veröffentlicht: (2024)
von: Alaofi, Marwah, et al.
Veröffentlicht: (2024)
ReFormeR: Learning and Applying Explicit Query Reformulation Patterns
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
Benchmarking LLM-based Relevance Judgment Methods
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
von: Amirshahi, Shakiba, et al.
Veröffentlicht: (2025)
von: Amirshahi, Shakiba, et al.
Veröffentlicht: (2025)
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2025)
QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025)
A Reproducibility Study of LLM-Based Query Reformulation
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
von: Arabzadeh, Ardalan
Veröffentlicht: (2025)
von: Arabzadeh, Ardalan
Veröffentlicht: (2025)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
von: Ebrahimi, Sajad, et al.
Veröffentlicht: (2025)
von: Ebrahimi, Sajad, et al.
Veröffentlicht: (2025)
Led to Mislead: Adversarial Content Injection for Attacks on Neural Ranking Models
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
von: Bigdeli, Amin, et al.
Veröffentlicht: (2026)
Benchmarking Prompt Sensitivity in Large Language Models
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
Predicting Retrieval Utility and Answer Quality in Retrieval-Augmented Generation
von: Tian, Fangzheng, et al.
Veröffentlicht: (2026)
von: Tian, Fangzheng, et al.
Veröffentlicht: (2026)
RAG over Thinking Traces Can Improve Reasoning Tasks
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
Annotative Indexing
von: Clarke, Charles L. A.
Veröffentlicht: (2024)
von: Clarke, Charles L. A.
Veröffentlicht: (2024)
Fishing for Answers: Exploring One-shot vs. Iterative Retrieval Strategies for Retrieval Augmented Generation
von: Lin, Huifeng, et al.
Veröffentlicht: (2025)
von: Lin, Huifeng, et al.
Veröffentlicht: (2025)
Ranking Generated Answers: On the Agreement of Retrieval Models with Humans on Consumer Health Questions
von: Heineking, Sebastian, et al.
Veröffentlicht: (2024)
von: Heineking, Sebastian, et al.
Veröffentlicht: (2024)
Answer Retrieval in Legal Community Question Answering
von: Askari, Arian, et al.
Veröffentlicht: (2024)
von: Askari, Arian, et al.
Veröffentlicht: (2024)
T$^2$-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation
von: Strich, Jan, et al.
Veröffentlicht: (2025)
von: Strich, Jan, et al.
Veröffentlicht: (2025)
Right Answer at the Right Time - Temporal Retrieval-Augmented Generation via Graph Summarization
von: Zhu, Zulun, et al.
Veröffentlicht: (2025)
von: Zhu, Zulun, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems
von: Caspari, Laura, et al.
Veröffentlicht: (2024)
von: Caspari, Laura, et al.
Veröffentlicht: (2024)
Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2024)
MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
von: Sun, Weiwei, et al.
Veröffentlicht: (2024)
von: Sun, Weiwei, et al.
Veröffentlicht: (2024)
Peerispect: Claim Verification in Scientific Peer Reviews
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval
von: Mao, Kelong, et al.
Veröffentlicht: (2024)
von: Mao, Kelong, et al.
Veröffentlicht: (2024)
ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
von: Ding, Hang, et al.
Veröffentlicht: (2025)
von: Ding, Hang, et al.
Veröffentlicht: (2025)
Beyond Utility: Evaluating LLM as Recommender
von: Jiang, Chumeng, et al.
Veröffentlicht: (2024)
von: Jiang, Chumeng, et al.
Veröffentlicht: (2024)
Green Recommender Systems: Optimizing Dataset Size for Energy-Efficient Algorithm Performance
von: Arabzadeh, Ardalan, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Ardalan, et al.
Veröffentlicht: (2024)
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
von: Zhang, Dake, et al.
Veröffentlicht: (2026)
von: Zhang, Dake, et al.
Veröffentlicht: (2026)
ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval
von: Choi, Hyewon, et al.
Veröffentlicht: (2026)
von: Choi, Hyewon, et al.
Veröffentlicht: (2026)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
von: Meng, Chuan, et al.
Veröffentlicht: (2024)
von: Meng, Chuan, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation Evaluation for Health Documents
von: Ceresa, Mario, et al.
Veröffentlicht: (2025)
von: Ceresa, Mario, et al.
Veröffentlicht: (2025)
LLM-based relevance assessment still can't replace human relevance assessment
von: Clarke, Charles L. A., et al.
Veröffentlicht: (2024)
von: Clarke, Charles L. A., et al.
Veröffentlicht: (2024)
WildClaims: Information Access Conversations in the Wild(Chat)
von: Joko, Hideaki, et al.
Veröffentlicht: (2025)
von: Joko, Hideaki, et al.
Veröffentlicht: (2025)
Interpretability Analysis of Domain Adapted Dense Retrievers
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
von: Bigdeli, Amin, et al.
Veröffentlicht: (2024) -
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024) -
Adversarial Attacks against Neural Ranking Models via In-Context Learning
von: Bigdeli, Amin, et al.
Veröffentlicht: (2025) -
A Comparison of Methods for Evaluating Generative IR
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024) -
Generative Information Retrieval Evaluation
von: Alaofi, Marwah, et al.
Veröffentlicht: (2024)