PyBench: Evaluating LLM Agent on various real-world coding tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yaolun, Pan, Yinxu, Wang, Yudong, Cai, Jie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
di: Shen, Haiyang, et al.
Pubblicazione: (2024)
di: Shen, Haiyang, et al.
Pubblicazione: (2024)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
di: He, Jiawei, et al.
Pubblicazione: (2026)
di: He, Jiawei, et al.
Pubblicazione: (2026)
DebugBench: Evaluating Debugging Capability of Large Language Models
di: Tian, Runchu, et al.
Pubblicazione: (2024)
di: Tian, Runchu, et al.
Pubblicazione: (2024)
Assessing LLM code generation quality through path planning tasks
di: Chen, Wanyi, et al.
Pubblicazione: (2025)
di: Chen, Wanyi, et al.
Pubblicazione: (2025)
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
di: Yin, Wenjing, et al.
Pubblicazione: (2025)
di: Yin, Wenjing, et al.
Pubblicazione: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
di: Qiu, Jielin, et al.
Pubblicazione: (2025)
di: Qiu, Jielin, et al.
Pubblicazione: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
di: Liu, Zhou, et al.
Pubblicazione: (2025)
di: Liu, Zhou, et al.
Pubblicazione: (2025)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
di: Jing, Lucas, et al.
Pubblicazione: (2026)
di: Jing, Lucas, et al.
Pubblicazione: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
di: Garg, Spandan, et al.
Pubblicazione: (2025)
di: Garg, Spandan, et al.
Pubblicazione: (2025)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
di: Kaliyev, Alibek T., et al.
Pubblicazione: (2026)
di: Kaliyev, Alibek T., et al.
Pubblicazione: (2026)
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
di: Wu, Yue, et al.
Pubblicazione: (2025)
di: Wu, Yue, et al.
Pubblicazione: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
di: Yan, Shuo, et al.
Pubblicazione: (2025)
di: Yan, Shuo, et al.
Pubblicazione: (2025)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
di: Wang, Yubang, et al.
Pubblicazione: (2026)
di: Wang, Yubang, et al.
Pubblicazione: (2026)
FullStack Bench: Evaluating LLMs as Full Stack Coders
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
di: Rank, Ben, et al.
Pubblicazione: (2026)
di: Rank, Ben, et al.
Pubblicazione: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
CrackMeBench: Binary Reverse Engineering for Agents
di: David, Isaac, et al.
Pubblicazione: (2026)
di: David, Isaac, et al.
Pubblicazione: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
di: Tu, Xinming, et al.
Pubblicazione: (2026)
di: Tu, Xinming, et al.
Pubblicazione: (2026)
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
di: Jia, Jin, et al.
Pubblicazione: (2026)
di: Jia, Jin, et al.
Pubblicazione: (2026)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
di: Xia, Boming, et al.
Pubblicazione: (2024)
di: Xia, Boming, et al.
Pubblicazione: (2024)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
di: Zhang, Zehua, et al.
Pubblicazione: (2025)
di: Zhang, Zehua, et al.
Pubblicazione: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
di: Zhang, Wentao, et al.
Pubblicazione: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
di: Zhang, Lei, et al.
Pubblicazione: (2026)
di: Zhang, Lei, et al.
Pubblicazione: (2026)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
di: Kogler, Leon, et al.
Pubblicazione: (2026)
di: Kogler, Leon, et al.
Pubblicazione: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
di: Zhou, Qixing, et al.
Pubblicazione: (2026)
di: Zhou, Qixing, et al.
Pubblicazione: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
di: Merrill, Mike A., et al.
Pubblicazione: (2026)
di: Merrill, Mike A., et al.
Pubblicazione: (2026)
Evolving Excellence: Automated Optimization of LLM-based Agents
di: Brookes, Paul, et al.
Pubblicazione: (2025)
di: Brookes, Paul, et al.
Pubblicazione: (2025)
Stop Comparing LLM Agents Without Disclosing the Harness
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
di: Zhang, Yunbei, et al.
Pubblicazione: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
di: Han, Tingxu, et al.
Pubblicazione: (2026)
di: Han, Tingxu, et al.
Pubblicazione: (2026)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
di: Zheng, Mingwei, et al.
Pubblicazione: (2025)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
di: Dong, Honghua, et al.
Pubblicazione: (2025)
di: Dong, Honghua, et al.
Pubblicazione: (2025)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
di: Tao, Wei, et al.
Pubblicazione: (2024)
di: Tao, Wei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
di: Shen, Haiyang, et al.
Pubblicazione: (2024) -
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
di: He, Jiawei, et al.
Pubblicazione: (2026) -
DebugBench: Evaluating Debugging Capability of Large Language Models
di: Tian, Runchu, et al.
Pubblicazione: (2024) -
Assessing LLM code generation quality through path planning tasks
di: Chen, Wanyi, et al.
Pubblicazione: (2025) -
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
di: Yin, Wenjing, et al.
Pubblicazione: (2025)