SlimRAG: Retrieval without Graphs via Entity-Aware Context Selection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jiale, Chen, Jiaxiang, Li, Zhucong, Ding, Jie, Zhao, Kui, Xu, Zenglin, Pang, Xin, Xu, Yinghui
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915353460736000
author Zhang, Jiale
Chen, Jiaxiang
Li, Zhucong
Ding, Jie
Zhao, Kui
Xu, Zenglin
Pang, Xin
Xu, Yinghui
author_facet Zhang, Jiale
Chen, Jiaxiang
Li, Zhucong
Ding, Jie
Zhao, Kui
Xu, Zenglin
Pang, Xin
Xu, Yinghui
contents Retrieval-Augmented Generation (RAG) enhances language models by incorporating external knowledge at inference time. However, graph-based RAG systems often suffer from structural overhead and imprecise retrieval: they require costly pipelines for entity linking and relation extraction, yet frequently return subgraphs filled with loosely related or tangential content. This stems from a fundamental flaw -- semantic similarity does not imply semantic relevance. We introduce SlimRAG, a lightweight framework for retrieval without graphs. SlimRAG replaces structure-heavy components with a simple yet effective entity-aware mechanism. At indexing time, it constructs a compact entity-to-chunk table based on semantic embeddings. At query time, it identifies salient entities, retrieves and scores associated chunks, and assembles a concise, contextually relevant input -- without graph traversal or edge construction. To quantify retrieval efficiency, we propose Relative Index Token Utilization (RITU), a metric measuring the compactness of retrieved content. Experiments across multiple QA benchmarks show that SlimRAG outperforms strong flat and graph-based baselines in accuracy while reducing index size and RITU (e.g., 16.31 vs. 56+), highlighting the value of structure-free, entity-centric context selection. The code will be released soon. https://github.com/continue-ai-company/SlimRAG
format Preprint
id arxiv_https___arxiv_org_abs_2506_17288
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SlimRAG: Retrieval without Graphs via Entity-Aware Context Selection
Zhang, Jiale
Chen, Jiaxiang
Li, Zhucong
Ding, Jie
Zhao, Kui
Xu, Zenglin
Pang, Xin
Xu, Yinghui
Information Retrieval
Artificial Intelligence
Computation and Language
Retrieval-Augmented Generation (RAG) enhances language models by incorporating external knowledge at inference time. However, graph-based RAG systems often suffer from structural overhead and imprecise retrieval: they require costly pipelines for entity linking and relation extraction, yet frequently return subgraphs filled with loosely related or tangential content. This stems from a fundamental flaw -- semantic similarity does not imply semantic relevance. We introduce SlimRAG, a lightweight framework for retrieval without graphs. SlimRAG replaces structure-heavy components with a simple yet effective entity-aware mechanism. At indexing time, it constructs a compact entity-to-chunk table based on semantic embeddings. At query time, it identifies salient entities, retrieves and scores associated chunks, and assembles a concise, contextually relevant input -- without graph traversal or edge construction. To quantify retrieval efficiency, we propose Relative Index Token Utilization (RITU), a metric measuring the compactness of retrieved content. Experiments across multiple QA benchmarks show that SlimRAG outperforms strong flat and graph-based baselines in accuracy while reducing index size and RITU (e.g., 16.31 vs. 56+), highlighting the value of structure-free, entity-centric context selection. The code will be released soon. https://github.com/continue-ai-company/SlimRAG
title SlimRAG: Retrieval without Graphs via Entity-Aware Context Selection
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.17288