Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hengran, Tang, Minghao, Bi, Keping, Guo, Jiafeng, Liu, Shihao, Shi, Daiting, Yin, Dawei, Cheng, Xueqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909831922712576
author Zhang, Hengran
Tang, Minghao
Bi, Keping
Guo, Jiafeng
Liu, Shihao
Shi, Daiting
Yin, Dawei
Cheng, Xueqi
author_facet Zhang, Hengran
Tang, Minghao
Bi, Keping
Guo, Jiafeng
Liu, Shihao
Shi, Daiting
Yin, Dawei
Cheng, Xueqi
contents This paper explores the use of large language models (LLMs) for annotating document utility in training retrieval and retrieval-augmented generation (RAG) systems, aiming to reduce dependence on costly human annotations. We address the gap between retrieval relevance and generative utility by employing LLMs to annotate document utility. To effectively utilize multiple positive samples per query, we introduce a novel loss that maximizes their summed marginal likelihood. Using the Qwen-2.5-32B model, we annotate utility on the MS MARCO dataset and conduct retrieval experiments on MS MARCO and BEIR, as well as RAG experiments on MS MARCO QA, NQ, and HotpotQA. Our results show that LLM-generated annotations enhance out-of-domain retrieval performance and improve RAG outcomes compared to models trained solely on human annotations or downstream QA metrics. Furthermore, combining LLM annotations with just 20% of human labels achieves performance comparable to using full human annotations. Our study offers a comprehensive approach to utilizing LLM annotations for initializing QA systems on new corpora.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
Zhang, Hengran
Tang, Minghao
Bi, Keping
Guo, Jiafeng
Liu, Shihao
Shi, Daiting
Yin, Dawei
Cheng, Xueqi
Information Retrieval
Artificial Intelligence
Computation and Language
This paper explores the use of large language models (LLMs) for annotating document utility in training retrieval and retrieval-augmented generation (RAG) systems, aiming to reduce dependence on costly human annotations. We address the gap between retrieval relevance and generative utility by employing LLMs to annotate document utility. To effectively utilize multiple positive samples per query, we introduce a novel loss that maximizes their summed marginal likelihood. Using the Qwen-2.5-32B model, we annotate utility on the MS MARCO dataset and conduct retrieval experiments on MS MARCO and BEIR, as well as RAG experiments on MS MARCO QA, NQ, and HotpotQA. Our results show that LLM-generated annotations enhance out-of-domain retrieval performance and improve RAG outcomes compared to models trained solely on human annotations or downstream QA metrics. Furthermore, combining LLM annotations with just 20% of human labels achieves performance comparable to using full human annotations. Our study offers a comprehensive approach to utilizing LLM annotations for initializing QA systems on new corpora.
title Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.05220