DependEval: Benchmarking LLMs for Repository Dependency Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Junjia, Liu, Yadi, Guo, Hongcheng, Wang, Jiawei, Huang, Haojian, Ni, Yunyi, Li, Zhoujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915189505392640
author Du, Junjia
Liu, Yadi
Guo, Hongcheng
Wang, Jiawei
Huang, Haojian
Ni, Yunyi
Li, Zhoujun
author_facet Du, Junjia
Liu, Yadi
Guo, Hongcheng
Wang, Jiawei
Huang, Haojian
Ni, Yunyi
Li, Zhoujun
contents While large language models (LLMs) have shown considerable promise in code generation, real-world software development demands advanced repository-level reasoning. This includes understanding dependencies, project structures, and managing multi-file changes. However, the ability of LLMs to effectively comprehend and handle complex code repositories has yet to be fully explored. To address challenges, we introduce a hierarchical benchmark designed to evaluate repository dependency understanding (DependEval). Benchmark is based on 15,576 repositories collected from real-world websites. It evaluates models on three core tasks: Dependency Recognition, Repository Construction, and Multi-file Editing, across 8 programming languages from actual code repositories. Our evaluation of over 25 LLMs reveals substantial performance gaps and provides valuable insights into repository-level code understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06689
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DependEval: Benchmarking LLMs for Repository Dependency Understanding
Du, Junjia
Liu, Yadi
Guo, Hongcheng
Wang, Jiawei
Huang, Haojian
Ni, Yunyi
Li, Zhoujun
Software Engineering
Computation and Language
While large language models (LLMs) have shown considerable promise in code generation, real-world software development demands advanced repository-level reasoning. This includes understanding dependencies, project structures, and managing multi-file changes. However, the ability of LLMs to effectively comprehend and handle complex code repositories has yet to be fully explored. To address challenges, we introduce a hierarchical benchmark designed to evaluate repository dependency understanding (DependEval). Benchmark is based on 15,576 repositories collected from real-world websites. It evaluates models on three core tasks: Dependency Recognition, Repository Construction, and Multi-file Editing, across 8 programming languages from actual code repositories. Our evaluation of over 25 LLMs reveals substantial performance gaps and provides valuable insights into repository-level code understanding.
title DependEval: Benchmarking LLMs for Repository Dependency Understanding
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2503.06689