A Benchmark for Localizing Code and Non-Code Issues in Software Projects
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zejun, Wang, Jian, Yang, Qingyun, Pan, Yifan, Tang, Yi, Li, Yi, Xing, Zhenchang, Zhang, Tian, Li, Xuandong, Zhang, Guoan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026)
by: Zhang, Zejun, et al.
Published: (2026)
Neurosymbolic Repo-level Code Localization
by: Xu, Xiufeng, et al.
Published: (2026)
by: Xu, Xiufeng, et al.
Published: (2026)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
In-Context Code-Text Learning for Bimodal Software Engineering
by: Tang, Xunzhu, et al.
Published: (2024)
by: Tang, Xunzhu, et al.
Published: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
CodeClash: Benchmarking Goal-Oriented Software Engineering
by: Yang, John, et al.
Published: (2025)
by: Yang, John, et al.
Published: (2025)
DevEval: Evaluating Code Generation in Practical Software Projects
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Code Review Agent Benchmark
by: Zhang, Yuntong, et al.
Published: (2026)
by: Zhang, Yuntong, et al.
Published: (2026)
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution
by: Li, KeFan, et al.
Published: (2025)
by: Li, KeFan, et al.
Published: (2025)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization
by: Reddy, Revanth Gangi, et al.
Published: (2025)
by: Reddy, Revanth Gangi, et al.
Published: (2025)
Multilingual Multimodal Software Developer for Code Generation
by: Chai, Linzheng, et al.
Published: (2025)
by: Chai, Linzheng, et al.
Published: (2025)
A^3-CodGen: A Repository-Level Code Generation Framework for Code Reuse with Local-Aware, Global-Aware, and Third-Party-Library-Aware
by: Liao, Dianshu, et al.
Published: (2023)
by: Liao, Dianshu, et al.
Published: (2023)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
by: Pan, Zhiyuan, et al.
Published: (2024)
by: Pan, Zhiyuan, et al.
Published: (2024)
SweRank: Software Issue Localization with Code Ranking
by: Reddy, Revanth Gangi, et al.
Published: (2025)
by: Reddy, Revanth Gangi, et al.
Published: (2025)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)
by: Cui, Yi
Published: (2025)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Survey of GenAI for Automotive Software Development: From Requirements to Executable Code
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language Models
by: Zhang, Zejun, et al.
Published: (2024)
by: Zhang, Zejun, et al.
Published: (2024)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
Adaptive Confidence Gating in Multi-Agent Collaboration for Efficient and Optimized Code Generation
by: Zhang, Haoji, et al.
Published: (2026)
by: Zhang, Haoji, et al.
Published: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
AI Code in the Wild: Measuring Security Risks and Ecosystem Shifts of AI-Generated Code in Modern Software
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
From Charts to Code: A Hierarchical Benchmark for Multimodal Models
by: Tang, Jiahao, et al.
Published: (2025)
by: Tang, Jiahao, et al.
Published: (2025)
Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension
by: Xu, Fangzhou, et al.
Published: (2024)
by: Xu, Fangzhou, et al.
Published: (2024)
Do Machines and Humans Focus on Similar Code? Exploring Explainability of Large Language Models in Code Summarization
by: Li, Jiliang, et al.
Published: (2024)
by: Li, Jiliang, et al.
Published: (2024)
Unveiling Project-Specific Bias in Neural Code Models
by: Li, Zhiming, et al.
Published: (2022)
by: Li, Zhiming, et al.
Published: (2022)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
Deep Learning for Code Intelligence: Survey, Benchmark and Toolkit
by: Wan, Yao, et al.
Published: (2023)
by: Wan, Yao, et al.
Published: (2023)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Turning the Tide: Repository-based Code Reflection
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
by: Liang, Qingyuan, et al.
Published: (2021)
by: Liang, Qingyuan, et al.
Published: (2021)
CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
by: Gao, Jun, et al.
Published: (2026)
by: Gao, Jun, et al.
Published: (2026)
A test-free semantic mistakes localization framework in Neural Code Translation
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Similar Items
-
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026) -
Neurosymbolic Repo-level Code Localization
by: Xu, Xiufeng, et al.
Published: (2026) -
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025) -
In-Context Code-Text Learning for Bimodal Software Engineering
by: Tang, Xunzhu, et al.
Published: (2024) -
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)