Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Qika, Zhao, Tianzhe, He, Kai, Peng, Zhen, Xu, Fangzhi, Huang, Ling, Ma, Jingying, Feng, Mengling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915129261555712
author Lin, Qika
Zhao, Tianzhe
He, Kai
Peng, Zhen
Xu, Fangzhi
Huang, Ling
Ma, Jingying
Feng, Mengling
author_facet Lin, Qika
Zhao, Tianzhe
He, Kai
Peng, Zhen
Xu, Fangzhi
Huang, Ling
Ma, Jingying
Feng, Mengling
contents Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural language, the effective integration of holistic structural information of KGs with Large Language Models (LLMs) has emerged as a significant question. To this end, we propose a two-stage framework to learn and apply quantized codes for each entity, aiming for the seamless integration of KGs with LLMs. Firstly, a self-supervised quantized representation (SSQR) method is proposed to compress both KG structural and semantic knowledge into discrete codes (\ie, tokens) that align the format of language sentences. We further design KG instruction-following data by viewing these learned codes as features to directly input to LLMs, thereby achieving seamless integration. The experiment results demonstrate that SSQR outperforms existing unsupervised quantized methods, producing more distinguishable codes. Further, the fine-tuned LLaMA2 and LLaMA3.1 also have superior performance on KG link prediction and triple classification tasks, utilizing only 16 tokens per entity instead of thousands in conventional prompting methods.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models
Lin, Qika
Zhao, Tianzhe
He, Kai
Peng, Zhen
Xu, Fangzhi
Huang, Ling
Ma, Jingying
Feng, Mengling
Computation and Language
Artificial Intelligence
Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural language, the effective integration of holistic structural information of KGs with Large Language Models (LLMs) has emerged as a significant question. To this end, we propose a two-stage framework to learn and apply quantized codes for each entity, aiming for the seamless integration of KGs with LLMs. Firstly, a self-supervised quantized representation (SSQR) method is proposed to compress both KG structural and semantic knowledge into discrete codes (\ie, tokens) that align the format of language sentences. We further design KG instruction-following data by viewing these learned codes as features to directly input to LLMs, thereby achieving seamless integration. The experiment results demonstrate that SSQR outperforms existing unsupervised quantized methods, producing more distinguishable codes. Further, the fine-tuned LLaMA2 and LLaMA3.1 also have superior performance on KG link prediction and triple classification tasks, utilizing only 16 tokens per entity instead of thousands in conventional prompting methods.
title Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.18119