Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yifan, Yu, Xiaoyan, Pan, Tengfei, Li, Angsheng, Du, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912370479071232
author Wei, Yifan
Yu, Xiaoyan
Pan, Tengfei
Li, Angsheng
Du, Li
author_facet Wei, Yifan
Yu, Xiaoyan
Pan, Tengfei
Li, Angsheng
Du, Li
contents Large language models (LLMs) have achieved unprecedented performance by leveraging vast pretraining corpora, yet their performance remains suboptimal in knowledge-intensive domains such as medicine and scientific research, where high factual precision is required. While synthetic data provides a promising avenue for augmenting domain knowledge, existing methods frequently generate redundant samples that do not align with the model's true knowledge gaps. To overcome this limitation, we propose a novel Structural Entropy-guided Knowledge Navigator (SENATOR) framework that addresses the intrinsic knowledge deficiencies of LLMs. Our approach employs the Structure Entropy (SE) metric to quantify uncertainty along knowledge graph paths and leverages Monte Carlo Tree Search (MCTS) to selectively explore regions where the model lacks domain-specific knowledge. Guided by these insights, the framework generates targeted synthetic data for supervised fine-tuning, enabling continuous self-improvement. Experimental results on LLaMA-3 and Qwen2 across multiple domain-specific benchmarks show that SENATOR effectively detects and repairs knowledge deficiencies, achieving notable performance improvements. The code and data for our methods and experiments are available at https://github.com/weiyifan1023/senator.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs
Wei, Yifan
Yu, Xiaoyan
Pan, Tengfei
Li, Angsheng
Du, Li
Computation and Language
Large language models (LLMs) have achieved unprecedented performance by leveraging vast pretraining corpora, yet their performance remains suboptimal in knowledge-intensive domains such as medicine and scientific research, where high factual precision is required. While synthetic data provides a promising avenue for augmenting domain knowledge, existing methods frequently generate redundant samples that do not align with the model's true knowledge gaps. To overcome this limitation, we propose a novel Structural Entropy-guided Knowledge Navigator (SENATOR) framework that addresses the intrinsic knowledge deficiencies of LLMs. Our approach employs the Structure Entropy (SE) metric to quantify uncertainty along knowledge graph paths and leverages Monte Carlo Tree Search (MCTS) to selectively explore regions where the model lacks domain-specific knowledge. Guided by these insights, the framework generates targeted synthetic data for supervised fine-tuning, enabling continuous self-improvement. Experimental results on LLaMA-3 and Qwen2 across multiple domain-specific benchmarks show that SENATOR effectively detects and repairs knowledge deficiencies, achieving notable performance improvements. The code and data for our methods and experiments are available at https://github.com/weiyifan1023/senator.
title Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs
topic Computation and Language
url https://arxiv.org/abs/2505.07184