FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Qixing, Zhang, Jiacheng, Wang, Haiyang, Hao, Rui, Wang, Jiahe, Han, Minghao, Yang, Yuxue, Wu, Shuzhe, Pan, Feiyang, Fan, Lue, Tu, Dandan, Zhang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
von: Lin, Yusong, et al.
Veröffentlicht: (2026)
von: Lin, Yusong, et al.
Veröffentlicht: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
von: Li, Haiyang
Veröffentlicht: (2025)
von: Li, Haiyang
Veröffentlicht: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
von: Yang, Yuxue, et al.
Veröffentlicht: (2024)
von: Yang, Yuxue, et al.
Veröffentlicht: (2024)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis
von: Wang, Yunkun, et al.
Veröffentlicht: (2026)
von: Wang, Yunkun, et al.
Veröffentlicht: (2026)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
von: Wang, Yilun, et al.
Veröffentlicht: (2026)
von: Wang, Yilun, et al.
Veröffentlicht: (2026)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
von: Xu, Xinbo, et al.
Veröffentlicht: (2026)
von: Xu, Xinbo, et al.
Veröffentlicht: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
von: Feng, Jia, et al.
Veröffentlicht: (2024)
von: Feng, Jia, et al.
Veröffentlicht: (2024)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
von: Wang, Yubang, et al.
Veröffentlicht: (2026)
von: Wang, Yubang, et al.
Veröffentlicht: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
von: Lin, Yusong, et al.
Veröffentlicht: (2026)
von: Lin, Yusong, et al.
Veröffentlicht: (2026)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
von: Feng, Yukang, et al.
Veröffentlicht: (2026)
von: Feng, Yukang, et al.
Veröffentlicht: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
von: Liu, Shuhan, et al.
Veröffentlicht: (2026)
von: Liu, Shuhan, et al.
Veröffentlicht: (2026)
RubberDuckBench: A Benchmark for AI Coding Assistants
von: Mohammed, Ferida, et al.
Veröffentlicht: (2026)
von: Mohammed, Ferida, et al.
Veröffentlicht: (2026)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
Program Feature-based Fuzzing Benchmarking
von: Miao, Miao
Veröffentlicht: (2025)
von: Miao, Miao
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
von: Zhang, Zejun, et al.
Veröffentlicht: (2025)
von: Zhang, Zejun, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
RepoSummary: Feature-Oriented Summarization and Documentation Generation for Code Repositories
von: Zhu, Yifeng, et al.
Veröffentlicht: (2025)
von: Zhu, Yifeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
von: Lin, Yusong, et al.
Veröffentlicht: (2026) -
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026) -
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025) -
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025) -
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
von: Li, Haiyang
Veröffentlicht: (2025)