Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vogel, Martin, Meyer-Eschenbach, Falk, Kohler, Severin, Grünewald, Elias, Balzer, Felix
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918414545584128
author Vogel, Martin
Meyer-Eschenbach, Falk
Kohler, Severin
Grünewald, Elias
Balzer, Felix
author_facet Vogel, Martin
Meyer-Eschenbach, Falk
Kohler, Severin
Grünewald, Elias
Balzer, Felix
contents Large Language Model (LLM) coding agents typically explore codebases through repeated file-reading and grep-searching, consuming thousands of tokens per query without structural understanding. We present Codebase-Memory, an open-source system that constructs a persistent, Tree-Sitter-based knowledge graph via the Model Context Protocol (MCP), parsing 66 languages through a multi-phase pipeline with parallel worker pools, call-graph traversal, impact analysis, and community discovery. Evaluated across 31 real-world repositories, Codebase-Memory achieves 83% answer quality versus 92% for a file-exploration agent, at ten times fewer tokens and 2.1 times fewer tool calls. For graph-native queries such as hub detection and caller ranking, it matches or exceeds the explorer on 19 of 31 languages.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27277
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP
Vogel, Martin
Meyer-Eschenbach, Falk
Kohler, Severin
Grünewald, Elias
Balzer, Felix
Software Engineering
Artificial Intelligence
Programming Languages
D.2.3; D.2.7; D.3.4; H.3.3; I.2.2
Large Language Model (LLM) coding agents typically explore codebases through repeated file-reading and grep-searching, consuming thousands of tokens per query without structural understanding. We present Codebase-Memory, an open-source system that constructs a persistent, Tree-Sitter-based knowledge graph via the Model Context Protocol (MCP), parsing 66 languages through a multi-phase pipeline with parallel worker pools, call-graph traversal, impact analysis, and community discovery. Evaluated across 31 real-world repositories, Codebase-Memory achieves 83% answer quality versus 92% for a file-exploration agent, at ten times fewer tokens and 2.1 times fewer tool calls. For graph-native queries such as hub detection and caller ranking, it matches or exceeds the explorer on 19 of 31 languages.
title Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP
topic Software Engineering
Artificial Intelligence
Programming Languages
D.2.3; D.2.7; D.3.4; H.3.3; I.2.2
url https://arxiv.org/abs/2603.27277