Towards Self-Evolving Agentic Literature Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Yuwen, Jin, Tian, Kang, Jing, Pang, Xianghe, Chai, Jingyi, Miao, Tingjia, Liu, Fenyi, Wang, WenHao, Yao, Sikai, Zhang, Yuzhi, Chen, Siheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910218940579840
author Du, Yuwen
Jin, Tian
Kang, Jing
Pang, Xianghe
Chai, Jingyi
Miao, Tingjia
Liu, Fenyi
Wang, WenHao
Yao, Sikai
Zhang, Yuzhi
Chen, Siheng
author_facet Du, Yuwen
Jin, Tian
Kang, Jing
Pang, Xianghe
Chai, Jingyi
Miao, Tingjia
Liu, Fenyi
Wang, WenHao
Yao, Sikai
Zhang, Yuzhi
Chen, Siheng
contents As large language models reshape scientific research, literature retrieval faces a twofold challenge: ensuring source authenticity while maintaining a deep comprehension of academic search intents. While reliable, traditional keyword-centric search fails to capture complex research intents. Frontier LLMs can handle complex research intents, but their high cost and tendency to hallucinate remain key limitations. Here we introduce PaSaMaster, a self-evolving agentic literature retrieval system that produces relevance-scored paper rankings with evidence-grounded recommendations through iterative intent analysis, retrieval, and ranking. It is built on three key designs. First, it transforms literature retrieval from a one shot query--document matching problem into a search process that evolves over time, using ranked evidence to reveal gaps, refine intents, and guide follow-up searches. Second, it prevents hallucinated sources by treating retrieval as intent--paper relevance ranking rather than generation. Finally, PaSaMaster improves cost efficiency by separating planning from retrieval: a frontier LLM is used only for intent understanding, while large scale retrieval and relevance scoring are delegated to customized corpora and lightweight models. Evaluated on the PaSaMaster Benchmark across 38 scientific disciplines, our system exposes the severe inaccuracy and incompleteness of traditional keyword retrieval (improving F1-score by 15.6X) and the unreliability of generative LLMs (which exhibit hallucination rates up to 37.79%). Remarkably, PaSaMaster outperforms GPT-5.2 by 30.0% at a mere 1% of the computational cost while ensuring zero source hallucination: https://github.com/sjtu-sai-agents/PaSaMaster
format Preprint
id arxiv_https___arxiv_org_abs_2605_14306
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Self-Evolving Agentic Literature Retrieval
Du, Yuwen
Jin, Tian
Kang, Jing
Pang, Xianghe
Chai, Jingyi
Miao, Tingjia
Liu, Fenyi
Wang, WenHao
Yao, Sikai
Zhang, Yuzhi
Chen, Siheng
Information Retrieval
As large language models reshape scientific research, literature retrieval faces a twofold challenge: ensuring source authenticity while maintaining a deep comprehension of academic search intents. While reliable, traditional keyword-centric search fails to capture complex research intents. Frontier LLMs can handle complex research intents, but their high cost and tendency to hallucinate remain key limitations. Here we introduce PaSaMaster, a self-evolving agentic literature retrieval system that produces relevance-scored paper rankings with evidence-grounded recommendations through iterative intent analysis, retrieval, and ranking. It is built on three key designs. First, it transforms literature retrieval from a one shot query--document matching problem into a search process that evolves over time, using ranked evidence to reveal gaps, refine intents, and guide follow-up searches. Second, it prevents hallucinated sources by treating retrieval as intent--paper relevance ranking rather than generation. Finally, PaSaMaster improves cost efficiency by separating planning from retrieval: a frontier LLM is used only for intent understanding, while large scale retrieval and relevance scoring are delegated to customized corpora and lightweight models. Evaluated on the PaSaMaster Benchmark across 38 scientific disciplines, our system exposes the severe inaccuracy and incompleteness of traditional keyword retrieval (improving F1-score by 15.6X) and the unreliability of generative LLMs (which exhibit hallucination rates up to 37.79%). Remarkably, PaSaMaster outperforms GPT-5.2 by 30.0% at a mere 1% of the computational cost while ensuring zero source hallucination: https://github.com/sjtu-sai-agents/PaSaMaster
title Towards Self-Evolving Agentic Literature Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2605.14306