TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Steven, Luo, Jane, Zhang, Xin, Liu, Aofan, Liu, Hao, Wu, Jie, Huang, Ziyang, Huang, Yangyu, Kang, Yu, Li, Scarlett |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Closing the Loop: Universal Repository Representation with RPG-Encoder
by: Luo, Jane, et al.
Published: (2026)
by: Luo, Jane, et al.
Published: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
BugForge: Constructing and Utilizing DBMS Bug Repository to Enhance DBMS Testing
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
by: Luo, Jane, et al.
Published: (2025)
by: Luo, Jane, et al.
Published: (2025)
Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
by: Liu, Kaibo, et al.
Published: (2024)
by: Liu, Kaibo, et al.
Published: (2024)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
by: Du, Junjia, et al.
Published: (2025)
by: Du, Junjia, et al.
Published: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
by: Xie, Yiqing, et al.
Published: (2025)
by: Xie, Yiqing, et al.
Published: (2025)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
by: Guo, Liwei, et al.
Published: (2025)
by: Guo, Liwei, et al.
Published: (2025)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
by: Liu, Linbo, et al.
Published: (2025)
by: Liu, Linbo, et al.
Published: (2025)
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
by: Peng, Zhiyuan, et al.
Published: (2025)
by: Peng, Zhiyuan, et al.
Published: (2025)
iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test Generation
by: Wang, Junyi, et al.
Published: (2026)
by: Wang, Junyi, et al.
Published: (2026)
Dependency-Guided Repository-Level C-to-Rust Translation with Reinforcement Alignment
by: Feng, Jia, et al.
Published: (2026)
by: Feng, Jia, et al.
Published: (2026)
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
by: Li, Kefan, et al.
Published: (2026)
by: Li, Kefan, et al.
Published: (2026)
How is Testing Related to Single Statement Bugs?
by: Rahman, Habibur, et al.
Published: (2024)
by: Rahman, Habibur, et al.
Published: (2024)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies
by: Widyasari, Ratnadira, et al.
Published: (2024)
by: Widyasari, Ratnadira, et al.
Published: (2024)
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
by: Wu, Qiushi, et al.
Published: (2025)
by: Wu, Qiushi, et al.
Published: (2025)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
by: Li, Xuefeng, et al.
Published: (2026)
by: Li, Xuefeng, et al.
Published: (2026)
Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization
by: Caumartin, Genevieve, et al.
Published: (2025)
by: Caumartin, Genevieve, et al.
Published: (2025)
GUIPilot: A Consistency-based Mobile GUI Testing Approach for Detecting Application-specific Bugs
by: Liu, Ruofan, et al.
Published: (2025)
by: Liu, Ruofan, et al.
Published: (2025)
Automated Duplicate Bug Report Detection in Large Open Bug Repositories
by: Laney, Clare E., et al.
Published: (2025)
by: Laney, Clare E., et al.
Published: (2025)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
by: Ou, Guangsheng, et al.
Published: (2024)
by: Ou, Guangsheng, et al.
Published: (2024)
Enriching Automatic Test Case Generation by Extracting Relevant Test Inputs from Bug Reports
by: Ouédraogo, Wendkûuni C., et al.
Published: (2023)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2023)
Fix the Tests: Augmenting LLMs to Repair Test Cases with Static Collector and Neural Reranker
by: Liu, Jun, et al.
Published: (2024)
by: Liu, Jun, et al.
Published: (2024)
Evolving Triple Knowledge-Augmented LLMs for Code Translation in Repository Context
by: Ou, Guangsheng, et al.
Published: (2025)
by: Ou, Guangsheng, et al.
Published: (2025)
A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Write Your Own CodeChecker: An Automated Test-Driven Checker Development Approach with LLMs
by: Liu, Jun, et al.
Published: (2024)
by: Liu, Jun, et al.
Published: (2024)
Understanding Bug-Reproducing Tests: A First Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
AOCI: Symbolic-Semantic Indexing for Practical Repository-Scale Code Understanding with LLMs
by: Liu, Jinshi, et al.
Published: (2026)
by: Liu, Jinshi, et al.
Published: (2026)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
by: Mang, Qiuyang, et al.
Published: (2025)
by: Mang, Qiuyang, et al.
Published: (2025)
ExploraCoder: Advancing code generation for multiple unseen APIs via planning and chained exploration
by: Wang, Yunkun, et al.
Published: (2024)
by: Wang, Yunkun, et al.
Published: (2024)
Finding XPath Bugs in XML Document Processors via Differential Testing
by: Li, Shuxin, et al.
Published: (2024)
by: Li, Shuxin, et al.
Published: (2024)
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
by: Du, Xueying, et al.
Published: (2026)
by: Du, Xueying, et al.
Published: (2026)
Isolating Compiler Bugs through Compilation Steps Analysis
by: Liu, Yujie, et al.
Published: (2025)
by: Liu, Yujie, et al.
Published: (2025)
Similar Items
-
Closing the Loop: Universal Repository Representation with RPG-Encoder
by: Luo, Jane, et al.
Published: (2026) -
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025) -
BugForge: Constructing and Utilizing DBMS Bug Repository to Enhance DBMS Testing
by: Li, Dawei, et al.
Published: (2026) -
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
by: Luo, Jane, et al.
Published: (2025) -
Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification
by: Liu, Aofan, et al.
Published: (2025)