SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Lu, Xu, Yijie, Ye, Jinhui, Liu, Hao, Xiong, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909915754266624
author Dai, Lu
Xu, Yijie
Ye, Jinhui
Liu, Hao
Xiong, Hui
author_facet Dai, Lu
Xu, Yijie
Ye, Jinhui
Liu, Hao
Xiong, Hui
contents Large Language Models (LLMs) have demonstrated improved generation performance by incorporating externally retrieved knowledge, a process known as retrieval-augmented generation (RAG). Despite the potential of this approach, existing studies evaluate RAG effectiveness by 1) assessing retrieval and generation components jointly, which obscures retrieval's distinct contribution, or 2) examining retrievers using traditional metrics such as NDCG, which creates a gap in understanding retrieval's true utility in the overall generation process. To address the above limitations, in this work, we introduce an automatic evaluation method that measures retrieval quality through the lens of information gain within the RAG framework. Specifically, we propose Semantic Perplexity (SePer), a metric that captures the LLM's internal belief about the correctness of the retrieved information. We quantify the utility of retrieval by the extent to which it reduces semantic perplexity post-retrieval. Extensive experiments demonstrate that SePer not only aligns closely with human preferences but also offers a more precise and efficient evaluation of retrieval utility across diverse RAG scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
Dai, Lu
Xu, Yijie
Ye, Jinhui
Liu, Hao
Xiong, Hui
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have demonstrated improved generation performance by incorporating externally retrieved knowledge, a process known as retrieval-augmented generation (RAG). Despite the potential of this approach, existing studies evaluate RAG effectiveness by 1) assessing retrieval and generation components jointly, which obscures retrieval's distinct contribution, or 2) examining retrievers using traditional metrics such as NDCG, which creates a gap in understanding retrieval's true utility in the overall generation process. To address the above limitations, in this work, we introduce an automatic evaluation method that measures retrieval quality through the lens of information gain within the RAG framework. Specifically, we propose Semantic Perplexity (SePer), a metric that captures the LLM's internal belief about the correctness of the retrieved information. We quantify the utility of retrieval by the extent to which it reduces semantic perplexity post-retrieval. Extensive experiments demonstrate that SePer not only aligns closely with human preferences but also offers a more precise and efficient evaluation of retrieval utility across diverse RAG scenarios.
title SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.01478