SEARA: An Automated Approach for Obtaining Optimal Retrievers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yuheng, Zou, Yiran, Wang, Yuzhu, Tian, Min, Zhu, Yanhua, Huang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916973968883712
author Yuheng, Zou
Yiran, Wang
Yuzhu, Tian
Min, Zhu
Yanhua, Huang
author_facet Yuheng, Zou
Yiran, Wang
Yuzhu, Tian
Min, Zhu
Yanhua, Huang
contents Retrieval-Augmented Generation (RAG) is a core approach for enhancing Large Language Models (LLMs), where the effectiveness of the retriever largely determines the overall response quality of RAG systems. Retrievers encompass a multitude of hyperparameters that significantly impact performance outcomes and demonstrate sensitivity to specific applications. Nevertheless, hyperparameter optimization entails prohibitively high computational expenses. Existing evaluation methods suffer from either prohibitive costs or disconnection from domain-specific scenarios. This paper proposes SEARA (Subset sampling Evaluation for Automatic Retriever Assessment), which addresses evaluation data challenges through subset sampling techniques and achieves robust automated retriever evaluation by minimal retrieval facts extraction and comprehensive retrieval metrics. Based on real user queries, this method enables fully automated retriever evaluation at low cost, thereby obtaining optimal retriever for specific business scenarios. We validate our method across classic RAG applications in rednote, including knowledge-based Q\&A system and retrieval-based travel assistant, successfully obtaining scenario-specific optimal retrievers.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06554
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEARA: An Automated Approach for Obtaining Optimal Retrievers
Yuheng, Zou
Yiran, Wang
Yuzhu, Tian
Min, Zhu
Yanhua, Huang
Information Retrieval
Retrieval-Augmented Generation (RAG) is a core approach for enhancing Large Language Models (LLMs), where the effectiveness of the retriever largely determines the overall response quality of RAG systems. Retrievers encompass a multitude of hyperparameters that significantly impact performance outcomes and demonstrate sensitivity to specific applications. Nevertheless, hyperparameter optimization entails prohibitively high computational expenses. Existing evaluation methods suffer from either prohibitive costs or disconnection from domain-specific scenarios. This paper proposes SEARA (Subset sampling Evaluation for Automatic Retriever Assessment), which addresses evaluation data challenges through subset sampling techniques and achieves robust automated retriever evaluation by minimal retrieval facts extraction and comprehensive retrieval metrics. Based on real user queries, this method enables fully automated retriever evaluation at low cost, thereby obtaining optimal retriever for specific business scenarios. We validate our method across classic RAG applications in rednote, including knowledge-based Q\&A system and retrieval-based travel assistant, successfully obtaining scenario-specific optimal retrievers.
title SEARA: An Automated Approach for Obtaining Optimal Retrievers
topic Information Retrieval
url https://arxiv.org/abs/2507.06554