AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Haoyu, Tsang, Hong Ting, Bai, Jiaxin, Peng, Xi, Zhang, Gong, Song, Yangqiu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917402465271808
author Huang, Haoyu
Tsang, Hong Ting
Bai, Jiaxin
Peng, Xi
Zhang, Gong
Song, Yangqiu
author_facet Huang, Haoyu
Tsang, Hong Ting
Bai, Jiaxin
Peng, Xi
Zhang, Gong
Song, Yangqiu
contents Retrieval-augmented generation (RAG) has shown some success in augmenting large language models (LLMs) with external knowledge. However, as a non-parametric knowledge integration paradigm for LLMs, RAG methods heavily rely on external retrieval modules and the retrieved textual context prior. Especially for very large scale knowledge augmentation, they would introduce substantial inference latency due to expensive searches and much longer relevant context. In this paper, we propose a parametric knowledge integration method, called \textbf{AtlasKV}, a scalable, effective, and general way to augment LLMs with billion-scale knowledge graphs (KGs) (e.g. 1B triples) using very little GPU memory cost (e.g. less than 20GB VRAM). In AtlasKV, we introduce KG2KV and HiKVP to integrate KG triples into LLMs at scale with sub-linear time and memory complexity. It maintains strong knowledge grounding and generalization performance using the LLMs' inherent attention mechanism, and requires no external retrievers, long context priors, or retraining when adapting to new knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17934
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
Huang, Haoyu
Tsang, Hong Ting
Bai, Jiaxin
Peng, Xi
Zhang, Gong
Song, Yangqiu
Computation and Language
Artificial Intelligence
Retrieval-augmented generation (RAG) has shown some success in augmenting large language models (LLMs) with external knowledge. However, as a non-parametric knowledge integration paradigm for LLMs, RAG methods heavily rely on external retrieval modules and the retrieved textual context prior. Especially for very large scale knowledge augmentation, they would introduce substantial inference latency due to expensive searches and much longer relevant context. In this paper, we propose a parametric knowledge integration method, called \textbf{AtlasKV}, a scalable, effective, and general way to augment LLMs with billion-scale knowledge graphs (KGs) (e.g. 1B triples) using very little GPU memory cost (e.g. less than 20GB VRAM). In AtlasKV, we introduce KG2KV and HiKVP to integrate KG triples into LLMs at scale with sub-linear time and memory complexity. It maintains strong knowledge grounding and generalization performance using the LLMs' inherent attention mechanism, and requires no external retrievers, long context priors, or retraining when adapting to new knowledge.
title AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.17934