RAVine: Reality-Aligned Evaluation for Agentic Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yilong, Long, Xiang, Zheng, Zhi, Gao, Jinhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912512852623360
author Xu, Yilong
Long, Xiang
Zheng, Zhi
Gao, Jinhua
author_facet Xu, Yilong
Long, Xiang
Zheng, Zhi
Gao, Jinhua
contents Agentic search, as a more autonomous and adaptive paradigm of retrieval augmentation, is driving the evolution of intelligent search systems. However, existing evaluation frameworks fail to align well with the goals of agentic search. First, the complex queries commonly used in current benchmarks often deviate from realistic user search scenarios. Second, prior approaches tend to introduce noise when extracting ground truth for end-to-end evaluations, leading to distorted assessments at a fine-grained level. Third, most current frameworks focus solely on the quality of final answers, neglecting the evaluation of the iterative process inherent to agentic search. To address these limitations, we propose RAVine -- a Reality-Aligned eValuation framework for agentic LLMs with search. RAVine targets multi-point queries and long-form answers that better reflect user intents, and introduces an attributable ground truth construction strategy to enhance the accuracy of fine-grained evaluation. Moreover, RAVine examines model's interaction with search tools throughout the iterative process, and accounts for factors of efficiency. We benchmark a series of models using RAVine and derive several insights, which we hope will contribute to advancing the development of agentic search systems. The code and datasets are available at https://github.com/SwordFaith/RAVine.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAVine: Reality-Aligned Evaluation for Agentic Search
Xu, Yilong
Long, Xiang
Zheng, Zhi
Gao, Jinhua
Computation and Language
Artificial Intelligence
Information Retrieval
Agentic search, as a more autonomous and adaptive paradigm of retrieval augmentation, is driving the evolution of intelligent search systems. However, existing evaluation frameworks fail to align well with the goals of agentic search. First, the complex queries commonly used in current benchmarks often deviate from realistic user search scenarios. Second, prior approaches tend to introduce noise when extracting ground truth for end-to-end evaluations, leading to distorted assessments at a fine-grained level. Third, most current frameworks focus solely on the quality of final answers, neglecting the evaluation of the iterative process inherent to agentic search. To address these limitations, we propose RAVine -- a Reality-Aligned eValuation framework for agentic LLMs with search. RAVine targets multi-point queries and long-form answers that better reflect user intents, and introduces an attributable ground truth construction strategy to enhance the accuracy of fine-grained evaluation. Moreover, RAVine examines model's interaction with search tools throughout the iterative process, and accounts for factors of efficiency. We benchmark a series of models using RAVine and derive several insights, which we hope will contribute to advancing the development of agentic search systems. The code and datasets are available at https://github.com/SwordFaith/RAVine.
title RAVine: Reality-Aligned Evaluation for Agentic Search
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2507.16725