Knowledge Homophily in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sahu, Utkarsh, Qi, Zhisheng, Halappanavar, Mahantesh, Lipka, Nedim, Rossi, Ryan A., Dernoncourt, Franck, Zhang, Yu, Ma, Yao, Wang, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915731980943360
author Sahu, Utkarsh
Qi, Zhisheng
Halappanavar, Mahantesh
Lipka, Nedim
Rossi, Ryan A.
Dernoncourt, Franck
Zhang, Yu
Ma, Yao
Wang, Yu
author_facet Sahu, Utkarsh
Qi, Zhisheng
Halappanavar, Mahantesh
Lipka, Nedim
Rossi, Ryan A.
Dernoncourt, Franck
Zhang, Yu
Ma, Yao
Wang, Yu
contents Large Language Models (LLMs) have been increasingly studied as neural knowledge bases for supporting knowledge-intensive applications such as question answering and fact checking. However, the structural organization of their knowledge remains unexplored. Inspired by cognitive neuroscience findings, such as semantic clustering and priming, where knowing one fact increases the likelihood of recalling related facts, we investigate an analogous knowledge homophily pattern in LLMs. To this end, we map LLM knowledge into a graph representation through knowledge checking at both the triplet and entity levels. After that, we analyze the knowledgeability relationship between an entity and its neighbors, discovering that LLMs tend to possess a similar level of knowledge about entities positioned closer in the graph. Motivated by this homophily principle, we propose a Graph Neural Network (GNN) regression model to estimate entity-level knowledgeability scores for triplets by leveraging their neighborhood scores. The predicted knowledgeability enables us to prioritize checking less well-known triplets, thereby maximizing knowledge coverage under the same labeling budget. This not only improves the efficiency of active labeling for fine-tuning to inject knowledge into LLMs but also enhances multi-hop path retrieval in reasoning-intensive question answering.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge Homophily in Large Language Models
Sahu, Utkarsh
Qi, Zhisheng
Halappanavar, Mahantesh
Lipka, Nedim
Rossi, Ryan A.
Dernoncourt, Franck
Zhang, Yu
Ma, Yao
Wang, Yu
Machine Learning
Artificial Intelligence
Computation and Language
Social and Information Networks
Large Language Models (LLMs) have been increasingly studied as neural knowledge bases for supporting knowledge-intensive applications such as question answering and fact checking. However, the structural organization of their knowledge remains unexplored. Inspired by cognitive neuroscience findings, such as semantic clustering and priming, where knowing one fact increases the likelihood of recalling related facts, we investigate an analogous knowledge homophily pattern in LLMs. To this end, we map LLM knowledge into a graph representation through knowledge checking at both the triplet and entity levels. After that, we analyze the knowledgeability relationship between an entity and its neighbors, discovering that LLMs tend to possess a similar level of knowledge about entities positioned closer in the graph. Motivated by this homophily principle, we propose a Graph Neural Network (GNN) regression model to estimate entity-level knowledgeability scores for triplets by leveraging their neighborhood scores. The predicted knowledgeability enables us to prioritize checking less well-known triplets, thereby maximizing knowledge coverage under the same labeling budget. This not only improves the efficiency of active labeling for fine-tuning to inject knowledge into LLMs but also enhances multi-hop path retrieval in reasoning-intensive question answering.
title Knowledge Homophily in Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2509.23773