PerfBench: Can Agents Resolve Real-World Performance Bugs?
Fuente:
arXiv
Saved in:
| Main Authors: | Garg, Spandan, Moghaddam, Roshanak Zilouchian, Sundaresan, Neel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot
by: Garg, Spandan, et al.
Published: (2023)
by: Garg, Spandan, et al.
Published: (2023)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025)
by: Liang, Shanchao, et al.
Published: (2025)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
by: Gautam, Dhruv, et al.
Published: (2025)
by: Gautam, Dhruv, et al.
Published: (2025)
AutoDev: Automated AI-Driven Development
by: Tufano, Michele, et al.
Published: (2024)
by: Tufano, Michele, et al.
Published: (2024)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
Do AI Models Dream of Faster Code? An Empirical Study on LLM-Proposed Performance Improvements in Real-World Software
by: Yi, Lirong, et al.
Published: (2025)
by: Yi, Lirong, et al.
Published: (2025)
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
Interpreting Performance Profiles with Deep Learning
by: Liu, Zhuoran
Published: (2025)
by: Liu, Zhuoran
Published: (2025)
Predicting Software Performance with Divide-and-Learn
by: Gong, Jingzhi, et al.
Published: (2023)
by: Gong, Jingzhi, et al.
Published: (2023)
Prompting for Performance: Exploring LLMs for Configuring Software
by: Spieker, Helge, et al.
Published: (2025)
by: Spieker, Helge, et al.
Published: (2025)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
by: Pham, Minh V. T., et al.
Published: (2025)
by: Pham, Minh V. T., et al.
Published: (2025)
Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis
by: Kumar, Punit, et al.
Published: (2025)
by: Kumar, Punit, et al.
Published: (2025)
Root Cause Localization for Microservice Systems in Cloud-edge Collaborative Environments
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
by: Rosas, Miguel Romero, et al.
Published: (2024)
by: Rosas, Miguel Romero, et al.
Published: (2024)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
by: Krupp, Lars, et al.
Published: (2026)
by: Krupp, Lars, et al.
Published: (2026)
Learning Performance-Improving Code Edits
by: Shypula, Alexander, et al.
Published: (2023)
by: Shypula, Alexander, et al.
Published: (2023)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
Predicting Configuration Performance in Multiple Environments with Sequential Meta-learning
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
by: Traini, Luca, et al.
Published: (2024)
by: Traini, Luca, et al.
Published: (2024)
PerfCoder: Large Language Models for Interpretable Code Performance Optimization
by: Yang, Jiuding, et al.
Published: (2025)
by: Yang, Jiuding, et al.
Published: (2025)
VecTrans: Enhancing Compiler Auto-Vectorization through LLM-Assisted Code Transformations
by: Zheng, Zhongchun, et al.
Published: (2025)
by: Zheng, Zhongchun, et al.
Published: (2025)
Investigating Execution-Aware Language Models for Code Optimization
by: Di Menna, Federico, et al.
Published: (2025)
by: Di Menna, Federico, et al.
Published: (2025)
Kevin: Multi-Turn RL for Generating CUDA Kernels
by: Baronio, Carlo, et al.
Published: (2025)
by: Baronio, Carlo, et al.
Published: (2025)
Lookup multivariate Kolmogorov-Arnold Networks
by: Pozdnyakov, Sergey, et al.
Published: (2025)
by: Pozdnyakov, Sergey, et al.
Published: (2025)
Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering
by: Navneet, Satyam Kumar, et al.
Published: (2025)
by: Navneet, Satyam Kumar, et al.
Published: (2025)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2023)
by: Rajput, Saurabhsingh, et al.
Published: (2023)
LLM-Vectorizer: LLM-based Verified Loop Vectorizer
by: Taneja, Jubi, et al.
Published: (2024)
by: Taneja, Jubi, et al.
Published: (2024)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Performance Prediction for Large Systems via Text-to-Text Regression
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Similar Items
-
RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot
by: Garg, Spandan, et al.
Published: (2023) -
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025) -
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
by: Gautam, Dhruv, et al.
Published: (2025) -
AutoDev: Automated AI-Driven Development
by: Tufano, Michele, et al.
Published: (2024) -
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)