Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Alebachew, Yoseph Berhanu, Leary, Hunter, Vaishampayan, Swanand, Brown, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Guided Exploration of Large-Scale Codebases
by: Alebachew, Yoseph Berhanu
Published: (2025)
by: Alebachew, Yoseph Berhanu
Published: (2025)
Automatic Bias Detection in Source Code Review
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
Are We on the Same Page? Examining Developer Perception Alignment in Open Source Code Reviews
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025)
Towards Evidence-Based Tech Hiring Pipelines
by: Brown, Chris, et al.
Published: (2025)
by: Brown, Chris, et al.
Published: (2025)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
by: Kabir, Azmain, et al.
Published: (2024)
by: Kabir, Azmain, et al.
Published: (2024)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs
by: Dhulshette, Nilesh, et al.
Published: (2025)
by: Dhulshette, Nilesh, et al.
Published: (2025)
Automated Snippet-Alignment Data Augmentation for Code Translation
by: Zhang, Zhiming, et al.
Published: (2025)
by: Zhang, Zhiming, et al.
Published: (2025)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
by: Lian, Keke, et al.
Published: (2025)
by: Lian, Keke, et al.
Published: (2025)
On the Impacts of Contexts on Repository-Level Code Generation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
Evaluating Repository-level Software Documentation via Question Answering and Feature-Driven Development
by: Wang, Xinchen, et al.
Published: (2026)
by: Wang, Xinchen, et al.
Published: (2026)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
by: Hu, Ruida, et al.
Published: (2024)
by: Hu, Ruida, et al.
Published: (2024)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
by: Pan, Zhiyuan, et al.
Published: (2024)
by: Pan, Zhiyuan, et al.
Published: (2024)
Uncovering Intention through LLM-Driven Code Snippet Description Generation
by: Nugroho, Yusuf Sulistyo, et al.
Published: (2025)
by: Nugroho, Yusuf Sulistyo, et al.
Published: (2025)
REPOFUSE: Repository-Level Code Completion with Fused Dual Context
by: Liang, Ming, et al.
Published: (2024)
by: Liang, Ming, et al.
Published: (2024)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
by: Xianpeng, et al.
Published: (2026)
by: Xianpeng, et al.
Published: (2026)
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
by: Pan, Ruwei, et al.
Published: (2026)
by: Pan, Ruwei, et al.
Published: (2026)
In Line with Context: Repository-Level Code Generation via Context Inlining
by: Hu, Chao, et al.
Published: (2026)
by: Hu, Chao, et al.
Published: (2026)
ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation
by: Hong, Jiseung, et al.
Published: (2026)
by: Hong, Jiseung, et al.
Published: (2026)
Toward Executable Repository-Level Code Generation via Environment Alignment
by: Pan, Ruwei, et al.
Published: (2026)
by: Pan, Ruwei, et al.
Published: (2026)
Automated Customization of LLMs for Enterprise Code Repositories Using Semantic Scopes
by: Finkler, Ulrich, et al.
Published: (2026)
by: Finkler, Ulrich, et al.
Published: (2026)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
RepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
by: Pan, Zhongqiang, et al.
Published: (2026)
by: Pan, Zhongqiang, et al.
Published: (2026)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
by: Jiang, Tianyue, et al.
Published: (2026)
by: Jiang, Tianyue, et al.
Published: (2026)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
by: Abedu, Samuel, et al.
Published: (2024)
by: Abedu, Samuel, et al.
Published: (2024)
RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion
by: Phan, Huy N., et al.
Published: (2024)
by: Phan, Huy N., et al.
Published: (2024)
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
by: Gurioli, Andrea, et al.
Published: (2026)
by: Gurioli, Andrea, et al.
Published: (2026)
RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation
by: Zhong, Sicheng, et al.
Published: (2025)
by: Zhong, Sicheng, et al.
Published: (2025)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
DialogAgent: An Auto-engagement Agent for Code Question Answering Data Production
by: Liang, Xiaoyun, et al.
Published: (2024)
by: Liang, Xiaoyun, et al.
Published: (2024)
Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
by: Zhang, Tanghaoran, et al.
Published: (2024)
by: Zhang, Tanghaoran, et al.
Published: (2024)
RepoReviewer: A Local-First Multi-Agent Architecture for Repository-Level Code Review
by: Zhang, Peng
Published: (2026)
by: Zhang, Peng
Published: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
by: Cai, Songcheng, et al.
Published: (2026)
by: Cai, Songcheng, et al.
Published: (2026)
Similar Items
-
AI-Guided Exploration of Large-Scale Codebases
by: Alebachew, Yoseph Berhanu
Published: (2025) -
Automatic Bias Detection in Source Code Review
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025) -
Are We on the Same Page? Examining Developer Perception Alignment in Open Source Code Reviews
by: Alebachew, Yoseph Berhanu, et al.
Published: (2025) -
Towards Evidence-Based Tech Hiring Pipelines
by: Brown, Chris, et al.
Published: (2025) -
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)