Reliable Evaluation Protocol for Low-Precision Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Kisu, Jang, Yoonna, Jang, Hwanseok, Choi, Kenneth, Augenstein, Isabelle, Lim, Heuiseok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expanding Computation Spaces of LLMs at Inference Time
by: Jang, Yoonna, et al.
Published: (2025)
by: Jang, Yoonna, et al.
Published: (2025)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
by: Park, Jeiyoon, et al.
Published: (2022)
by: Park, Jeiyoon, et al.
Published: (2022)
Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations
by: Jang, Yoonna, et al.
Published: (2024)
by: Jang, Yoonna, et al.
Published: (2024)
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
by: Augenstein, Isabelle
Published: (2026)
by: Augenstein, Isabelle
Published: (2026)
Llamion Technical Report
by: Yang, Kisu, et al.
Published: (2026)
by: Yang, Kisu, et al.
Published: (2026)
RARe: Retrieval Augmented Retrieval with In-Context Examples
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
Improving Health Question Answering with Reliable and Time-Aware Evidence Retrieval
by: Vladika, Juraj, et al.
Published: (2024)
by: Vladika, Juraj, et al.
Published: (2024)
CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
by: Lee, Seungyoon, et al.
Published: (2026)
by: Lee, Seungyoon, et al.
Published: (2026)
MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal
by: Jang, Youngjoon, et al.
Published: (2026)
by: Jang, Youngjoon, et al.
Published: (2026)
EncouRAGe: Evaluating RAG Local, Fast, and Reliable
by: Strich, Jan, et al.
Published: (2025)
by: Strich, Jan, et al.
Published: (2025)
Resolving the Robustness-Precision Trade-off in Financial RAG through Hybrid Document-Routed Retrieval
by: Cheng, Zhiyuan, et al.
Published: (2026)
by: Cheng, Zhiyuan, et al.
Published: (2026)
LANGSAE EDITING: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal
by: Kim, Dongjun, et al.
Published: (2026)
by: Kim, Dongjun, et al.
Published: (2026)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
by: Shin, Joongmin, et al.
Published: (2026)
by: Shin, Joongmin, et al.
Published: (2026)
AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis
by: Liao, Callie C., et al.
Published: (2025)
by: Liao, Callie C., et al.
Published: (2025)
Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
by: Hu, Chuxuan, et al.
Published: (2025)
by: Hu, Chuxuan, et al.
Published: (2025)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
by: Choi, Chanyeol, et al.
Published: (2025)
by: Choi, Chanyeol, et al.
Published: (2025)
Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
by: Hilmi, Muhammad Anis Al, et al.
Published: (2026)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
by: Merola, Carlo, et al.
Published: (2025)
by: Merola, Carlo, et al.
Published: (2025)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
by: Khanghah, Kiarash Naghavi, et al.
Published: (2026)
by: Khanghah, Kiarash Naghavi, et al.
Published: (2026)
Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
by: Chun, Yongchan, et al.
Published: (2025)
by: Chun, Yongchan, et al.
Published: (2025)
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
by: Kim, Taehee, et al.
Published: (2026)
by: Kim, Taehee, et al.
Published: (2026)
Self-Describing Structured Data with Dual-Layer Guidance: A Lightweight Alternative to RAG for Precision Retrieval in Large-Scale LLM Knowledge Navigation
by: Liu, Hung Ming
Published: (2026)
by: Liu, Hung Ming
Published: (2026)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
Model-Document Protocol for AI Search
by: Qian, Hongjin, et al.
Published: (2025)
by: Qian, Hongjin, et al.
Published: (2025)
EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
by: Masayoshi, Kanato, et al.
Published: (2025)
by: Masayoshi, Kanato, et al.
Published: (2025)
Chunking, Retrieval, and Re-ranking: An Empirical Evaluation of RAG Architectures for Policy Document Question Answering
by: Maharjan, Anuj, et al.
Published: (2026)
by: Maharjan, Anuj, et al.
Published: (2026)
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
by: Brehme, Lorenz, et al.
Published: (2026)
by: Brehme, Lorenz, et al.
Published: (2026)
Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
by: Liu, Geng, et al.
Published: (2025)
by: Liu, Geng, et al.
Published: (2025)
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
by: Killingback, Julian, et al.
Published: (2025)
by: Killingback, Julian, et al.
Published: (2025)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
by: Hsu, Tz-Huan, et al.
Published: (2026)
by: Hsu, Tz-Huan, et al.
Published: (2026)
Similar Items
-
Expanding Computation Spaces of LLMs at Inference Time
by: Jang, Yoonna, et al.
Published: (2025) -
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
by: Park, Jeiyoon, et al.
Published: (2022) -
Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations
by: Jang, Yoonna, et al.
Published: (2024) -
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
by: Augenstein, Isabelle
Published: (2026) -
Llamion Technical Report
by: Yang, Kisu, et al.
Published: (2026)