Redefining Retrieval Evaluation in the Era of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Trappolini, Giovanni, Cuconasu, Florin, Filice, Simone, Maarek, Yoelle, Silvestri, Fabrizio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Power of Noise: Redefining Retrieval for RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024)
by: Cuconasu, Florin, et al.
Published: (2024)
Do RAG Systems Really Suffer From Positional Bias?
by: Cuconasu, Florin, et al.
Published: (2025)
by: Cuconasu, Florin, et al.
Published: (2025)
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024)
by: Cuconasu, Florin, et al.
Published: (2024)
RRAML: Reinforced Retrieval Augmented Machine Learning
by: Bacciu, Andrea, et al.
Published: (2023)
by: Bacciu, Andrea, et al.
Published: (2023)
The Distracting Effect: Understanding Irrelevant Passages in RAG
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana
by: Filice, Simone, et al.
Published: (2025)
by: Filice, Simone, et al.
Published: (2025)
LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
by: Carmel, David, et al.
Published: (2025)
by: Carmel, David, et al.
Published: (2025)
SIGIR 2025 -- LiveRAG Challenge Report
by: Carmel, David, et al.
Published: (2025)
by: Carmel, David, et al.
Published: (2025)
Multimodal Neural Databases
by: Trappolini, Giovanni, et al.
Published: (2023)
by: Trappolini, Giovanni, et al.
Published: (2023)
ECLIPSE: Contrastive Dimension Importance Estimation with Pseudo-Irrelevance Feedback for Dense Retrieval
by: D'Erasmo, Giulio, et al.
Published: (2024)
by: D'Erasmo, Giulio, et al.
Published: (2024)
Large Search Model: Redefining Search Stack in the Era of LLMs
by: Wang, Liang, et al.
Published: (2023)
by: Wang, Liang, et al.
Published: (2023)
A Reproducible Analysis of Sequential Recommender Systems
by: Betello, Filippo, et al.
Published: (2024)
by: Betello, Filippo, et al.
Published: (2024)
When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively
by: Labruna, Tiziano, et al.
Published: (2024)
by: Labruna, Tiziano, et al.
Published: (2024)
LTRR: Learning To Rank Retrievers for LLMs
by: Kim, To Eun, et al.
Published: (2025)
by: Kim, To Eun, et al.
Published: (2025)
Evaluating Retrieval Quality in Retrieval-Augmented Generation
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
by: Wang, Xiaolei, et al.
Published: (2023)
by: Wang, Xiaolei, et al.
Published: (2023)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)
by: Gienapp, Lukas, et al.
Published: (2023)
Open-World Evaluation for Retrieving Diverse Perspectives
by: Chen, Hung-Ting, et al.
Published: (2024)
by: Chen, Hung-Ting, et al.
Published: (2024)
Recommender Systems in the Era of Large Language Models (LLMs)
by: Zhao, Zihuai, et al.
Published: (2023)
by: Zhao, Zihuai, et al.
Published: (2023)
Evaluating Large Language Models for Cross-Lingual Retrieval
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
Building Russian Benchmark for Evaluation of Information Retrieval Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
by: Zhao, Xinran, et al.
Published: (2024)
by: Zhao, Xinran, et al.
Published: (2024)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs
by: Ge, Ziyu, et al.
Published: (2025)
by: Ge, Ziyu, et al.
Published: (2025)
Grounding Arabic LLMs in the Doha Historical Dictionary: Retrieval-Augmented Understanding of Quran and Hadith
by: Eltanbouly, Somaya, et al.
Published: (2026)
by: Eltanbouly, Somaya, et al.
Published: (2026)
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users
by: Hu, Mengxuan, et al.
Published: (2024)
by: Hu, Mengxuan, et al.
Published: (2024)
Evaluating LLMs for Gender Disparities in Notable Persons
by: Rhue, Lauren, et al.
Published: (2024)
by: Rhue, Lauren, et al.
Published: (2024)
Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
by: Chitale, Pranjal A., et al.
Published: (2025)
by: Chitale, Pranjal A., et al.
Published: (2025)
Evaluating the Retrieval Component in LLM-Based Question Answering Systems
by: Alinejad, Ashkan, et al.
Published: (2024)
by: Alinejad, Ashkan, et al.
Published: (2024)
Evaluating the Performance of LLMs on Technical Language Processing tasks
by: Kernycky, Andrew, et al.
Published: (2024)
by: Kernycky, Andrew, et al.
Published: (2024)
HEISIR: Hierarchical Expansion of Inverted Semantic Indexing for Training-free Retrieval of Conversational Data using LLMs
by: Kim, Sangyeop, et al.
Published: (2025)
by: Kim, Sangyeop, et al.
Published: (2025)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
by: Amirshahi, Shakiba, et al.
Published: (2025)
by: Amirshahi, Shakiba, et al.
Published: (2025)
Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
by: Liu, Qijiong, et al.
Published: (2025)
by: Liu, Qijiong, et al.
Published: (2025)
Text2Cypher Across Languages: Evaluating and Finetuning LLMs
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
Similar Items
-
The Power of Noise: Redefining Retrieval for RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024) -
Do RAG Systems Really Suffer From Positional Bias?
by: Cuconasu, Florin, et al.
Published: (2025) -
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024) -
RRAML: Reinforced Retrieval Augmented Machine Learning
by: Bacciu, Andrea, et al.
Published: (2023) -
The Distracting Effect: Understanding Irrelevant Passages in RAG
by: Amiraz, Chen, et al.
Published: (2025)