ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Wenbin, Wang, Xin, Chen, Jiaoyan, Guo, Lingbing, Li, Zhao, Chen, Zirui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917004459376640
author Guo, Wenbin
Wang, Xin
Chen, Jiaoyan
Guo, Lingbing
Li, Zhao
Chen, Zirui
author_facet Guo, Wenbin
Wang, Xin
Chen, Jiaoyan
Guo, Lingbing
Li, Zhao
Chen, Zirui
contents Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and generalization capabilities beyond traditional embedding-based approaches. However, existing LLM-based methods often struggle to fully exploit structured semantic representations, as the continuous embedding space of pretrained KG models is fundamentally misaligned with the discrete token space of LLMs. This discrepancy hinders effective semantic transfer and limits their performance. To address this challenge, we propose ReaLM, a novel and effective framework that bridges the gap between KG embeddings and LLM tokenization through the mechanism of residual vector quantization. ReaLM discretizes pretrained KG embeddings into compact code sequences and integrates them as learnable tokens within the LLM vocabulary, enabling seamless fusion of symbolic and contextual knowledge. Furthermore, we incorporate ontology-guided class constraints to enforce semantic consistency, refining entity predictions based on class-level compatibility. Extensive experiments on two widely used benchmark datasets demonstrate that ReaLM achieves state-of-the-art performance, confirming its effectiveness in aligning structured knowledge with large-scale language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models
Guo, Wenbin
Wang, Xin
Chen, Jiaoyan
Guo, Lingbing
Li, Zhao
Chen, Zirui
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and generalization capabilities beyond traditional embedding-based approaches. However, existing LLM-based methods often struggle to fully exploit structured semantic representations, as the continuous embedding space of pretrained KG models is fundamentally misaligned with the discrete token space of LLMs. This discrepancy hinders effective semantic transfer and limits their performance. To address this challenge, we propose ReaLM, a novel and effective framework that bridges the gap between KG embeddings and LLM tokenization through the mechanism of residual vector quantization. ReaLM discretizes pretrained KG embeddings into compact code sequences and integrates them as learnable tokens within the LLM vocabulary, enabling seamless fusion of symbolic and contextual knowledge. Furthermore, we incorporate ontology-guided class constraints to enforce semantic consistency, refining entity predictions based on class-level compatibility. Extensive experiments on two widely used benchmark datasets demonstrate that ReaLM achieves state-of-the-art performance, confirming its effectiveness in aligning structured knowledge with large-scale language models.
title ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.09711