Saved in:
Bibliographic Details
Main Authors: Zhang, Duzhen, Wang, Zixiao, Li, Zhong-Zhi, Yu, Yahan, Jia, Shuncheng, Dong, Jiahua, Xu, Haotian, Wu, Xing, Zhang, Yingying, Zhang, Tielin, Yang, Jie, Chen, Xiuying, Song, Le
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.12393
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909742602911744
author Zhang, Duzhen
Wang, Zixiao
Li, Zhong-Zhi
Yu, Yahan
Jia, Shuncheng
Dong, Jiahua
Xu, Haotian
Wu, Xing
Zhang, Yingying
Zhang, Tielin
Yang, Jie
Chen, Xiuying
Song, Le
author_facet Zhang, Duzhen
Wang, Zixiao
Li, Zhong-Zhi
Yu, Yahan
Jia, Shuncheng
Dong, Jiahua
Xu, Haotian
Wu, Xing
Zhang, Yingying
Zhang, Tielin
Yang, Jie
Chen, Xiuying
Song, Le
contents The rapid expansion of medical literature presents growing challenges for structuring and integrating domain knowledge at scale. Knowledge Graphs (KGs) offer a promising solution by enabling efficient retrieval, automated reasoning, and knowledge discovery. However, current KG construction methods often rely on supervised pipelines with limited generalizability or naively aggregate outputs from Large Language Models (LLMs), treating biomedical corpora as static and ignoring the temporal dynamics and contextual uncertainty of evolving knowledge. To address these limitations, we introduce MedKGent, a LLM agent framework for constructing temporally evolving medical KGs. Leveraging over 10 million PubMed abstracts published between 1975 and 2023, we simulate the emergence of biomedical knowledge via a fine-grained daily time series. MedKGent incrementally builds the KG in a day-by-day manner using two specialized agents powered by the Qwen2.5-32B-Instruct model. The Extractor Agent identifies knowledge triples and assigns confidence scores via sampling-based estimation, which are used to filter low-confidence extractions and inform downstream processing. The Constructor Agent incrementally integrates the retained triples into a temporally evolving graph, guided by confidence scores and timestamps to reinforce recurring knowledge and resolve conflicts. The resulting KG contains 156,275 entities and 2,971,384 relational triples. Quality assessments by two SOTA LLMs and three domain experts demonstrate an accuracy approaching 90%, with strong inter-rater agreement. To evaluate downstream utility, we conduct RAG across seven medical question answering benchmarks using five leading LLMs, consistently observing significant improvements over non-augmented baselines. Case studies further demonstrate the KG's value in literature-based drug repurposing via confidence-aware causal inference.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12393
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
Zhang, Duzhen
Wang, Zixiao
Li, Zhong-Zhi
Yu, Yahan
Jia, Shuncheng
Dong, Jiahua
Xu, Haotian
Wu, Xing
Zhang, Yingying
Zhang, Tielin
Yang, Jie
Chen, Xiuying
Song, Le
Computation and Language
Artificial Intelligence
The rapid expansion of medical literature presents growing challenges for structuring and integrating domain knowledge at scale. Knowledge Graphs (KGs) offer a promising solution by enabling efficient retrieval, automated reasoning, and knowledge discovery. However, current KG construction methods often rely on supervised pipelines with limited generalizability or naively aggregate outputs from Large Language Models (LLMs), treating biomedical corpora as static and ignoring the temporal dynamics and contextual uncertainty of evolving knowledge. To address these limitations, we introduce MedKGent, a LLM agent framework for constructing temporally evolving medical KGs. Leveraging over 10 million PubMed abstracts published between 1975 and 2023, we simulate the emergence of biomedical knowledge via a fine-grained daily time series. MedKGent incrementally builds the KG in a day-by-day manner using two specialized agents powered by the Qwen2.5-32B-Instruct model. The Extractor Agent identifies knowledge triples and assigns confidence scores via sampling-based estimation, which are used to filter low-confidence extractions and inform downstream processing. The Constructor Agent incrementally integrates the retained triples into a temporally evolving graph, guided by confidence scores and timestamps to reinforce recurring knowledge and resolve conflicts. The resulting KG contains 156,275 entities and 2,971,384 relational triples. Quality assessments by two SOTA LLMs and three domain experts demonstrate an accuracy approaching 90%, with strong inter-rater agreement. To evaluate downstream utility, we conduct RAG across seven medical question answering benchmarks using five leading LLMs, consistently observing significant improvements over non-augmented baselines. Case studies further demonstrate the KG's value in literature-based drug repurposing via confidence-aware causal inference.
title MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.12393