DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Minghao, Zeng, Qingcheng, Zhao, Xujiang, Liu, Yanchi, Yu, Wenchao, Du, Mengnan, Chen, Haifeng, Cheng, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911474967904256
author Guo, Minghao
Zeng, Qingcheng
Zhao, Xujiang
Liu, Yanchi
Yu, Wenchao
Du, Mengnan
Chen, Haifeng
Cheng, Wei
author_facet Guo, Minghao
Zeng, Qingcheng
Zhao, Xujiang
Liu, Yanchi
Yu, Wenchao
Du, Mengnan
Chen, Haifeng
Cheng, Wei
contents Large Language Models (LLMs) excel at many reasoning tasks but struggle with knowledge-intensive queries due to their inability to dynamically access up-to-date or domain-specific information. Retrieval-Augmented Generation (RAG) has emerged as a promising solution, enabling LLMs to ground their responses in external sources. However, existing RAG methods lack fine-grained control over both the query and source sides, often resulting in noisy retrieval and shallow reasoning. In this work, we introduce DeepSieve, an agentic RAG framework that incorporates information sieving via LLM-as-a-knowledge-router. DeepSieve decomposes complex queries into structured sub-questions and recursively routes each to the most suitable knowledge source, filtering irrelevant information through a multi-stage distillation process. Our design emphasizes modularity, transparency, and adaptability, leveraging recent advances in agentic system design. Experiments on multi-hop QA tasks across heterogeneous sources demonstrate improved reasoning depth, retrieval precision, and interpretability over conventional RAG approaches. Our codes are available at https://github.com/MinghoKwok/DeepSieve.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22050
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
Guo, Minghao
Zeng, Qingcheng
Zhao, Xujiang
Liu, Yanchi
Yu, Wenchao
Du, Mengnan
Chen, Haifeng
Cheng, Wei
Computation and Language
Large Language Models (LLMs) excel at many reasoning tasks but struggle with knowledge-intensive queries due to their inability to dynamically access up-to-date or domain-specific information. Retrieval-Augmented Generation (RAG) has emerged as a promising solution, enabling LLMs to ground their responses in external sources. However, existing RAG methods lack fine-grained control over both the query and source sides, often resulting in noisy retrieval and shallow reasoning. In this work, we introduce DeepSieve, an agentic RAG framework that incorporates information sieving via LLM-as-a-knowledge-router. DeepSieve decomposes complex queries into structured sub-questions and recursively routes each to the most suitable knowledge source, filtering irrelevant information through a multi-stage distillation process. Our design emphasizes modularity, transparency, and adaptability, leveraging recent advances in agentic system design. Experiments on multi-hop QA tasks across heterogeneous sources demonstrate improved reasoning depth, retrieval precision, and interpretability over conventional RAG approaches. Our codes are available at https://github.com/MinghoKwok/DeepSieve.
title DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
topic Computation and Language
url https://arxiv.org/abs/2507.22050