Comprehensive Evaluation of Large Language Models on Software Engineering Tasks: A Multi-Task Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Gunawan, Go Frendi, Amien, Mukhlis |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OODEval: Evaluating Large Language Models on Object-Oriented Design
di: Xiao, Bingxu, et al.
Pubblicazione: (2026)
di: Xiao, Bingxu, et al.
Pubblicazione: (2026)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
di: Alshaikh, Moaath, et al.
Pubblicazione: (2026)
di: Alshaikh, Moaath, et al.
Pubblicazione: (2026)
Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
di: Shah, Syed Muhammad Ashhar, et al.
Pubblicazione: (2026)
di: Shah, Syed Muhammad Ashhar, et al.
Pubblicazione: (2026)
How Quickly Do Development Teams Update Their Vulnerable Dependencies?
di: Rahman, Imranur, et al.
Pubblicazione: (2024)
di: Rahman, Imranur, et al.
Pubblicazione: (2024)
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
di: Paiva, Camila A., et al.
Pubblicazione: (2026)
di: Paiva, Camila A., et al.
Pubblicazione: (2026)
Source Code Hotspots: A Diagnostic Method for Quality Issues
di: Muzammil, Saleha, et al.
Pubblicazione: (2026)
di: Muzammil, Saleha, et al.
Pubblicazione: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
di: Iscan, Mehmet
Pubblicazione: (2026)
di: Iscan, Mehmet
Pubblicazione: (2026)
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
di: Li, Caihua, et al.
Pubblicazione: (2026)
di: Li, Caihua, et al.
Pubblicazione: (2026)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
di: Palacios, Diego Cabezas
Pubblicazione: (2026)
di: Palacios, Diego Cabezas
Pubblicazione: (2026)
GEML: A Grammar-based Evolutionary Machine Learning Approach for Design-Pattern Detection
di: Barbudo, Rafael, et al.
Pubblicazione: (2024)
di: Barbudo, Rafael, et al.
Pubblicazione: (2024)
Making Software Metrics Useful
di: Tempero, Ewan, et al.
Pubblicazione: (2026)
di: Tempero, Ewan, et al.
Pubblicazione: (2026)
Stabilization Without Simplification: A Two-Dimensional Model of Software Evolution
di: Furukawa, Masaru
Pubblicazione: (2026)
di: Furukawa, Masaru
Pubblicazione: (2026)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
di: Ashrafi, Nazmus
Pubblicazione: (2026)
di: Ashrafi, Nazmus
Pubblicazione: (2026)
Overhead Measurement Noise in Different Runtime Environments
di: Reichelt, David Georg, et al.
Pubblicazione: (2024)
di: Reichelt, David Georg, et al.
Pubblicazione: (2024)
Energy-Aware Decision Making in Software Stack Upgrades
di: Stocker, Mirko, et al.
Pubblicazione: (2026)
di: Stocker, Mirko, et al.
Pubblicazione: (2026)
Site Reliability Engineering (SRE) and Observations on SRE Process to Make Tasks Easier
di: Puli, Balaram
Pubblicazione: (2025)
di: Puli, Balaram
Pubblicazione: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
di: Weng, Haojun, et al.
Pubblicazione: (2026)
di: Weng, Haojun, et al.
Pubblicazione: (2026)
Using LLMs to Establish Implicit User Sentiment of Software Desirability
di: Weitl-Harms, Sherri, et al.
Pubblicazione: (2024)
di: Weitl-Harms, Sherri, et al.
Pubblicazione: (2024)
SPViz: A DSL-Driven Approach for Software Project Visualization Tooling
di: Rentz, Niklas, et al.
Pubblicazione: (2024)
di: Rentz, Niklas, et al.
Pubblicazione: (2024)
From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks
di: Weerasinghe, Mineth, et al.
Pubblicazione: (2026)
di: Weerasinghe, Mineth, et al.
Pubblicazione: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
di: Zadenoori, Mohammad Amin, et al.
Pubblicazione: (2025)
di: Zadenoori, Mohammad Amin, et al.
Pubblicazione: (2025)
Providing Information About Implemented Algorithms Improves Program Comprehension: A Controlled Experiment
di: Neumüller, Denis, et al.
Pubblicazione: (2025)
di: Neumüller, Denis, et al.
Pubblicazione: (2025)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
di: Mazaheri, Parsa
Pubblicazione: (2026)
di: Mazaheri, Parsa
Pubblicazione: (2026)
Evaluating Software Contribution Quality: Time-to-Modification Theory
di: Bishop III, Vincil, et al.
Pubblicazione: (2024)
di: Bishop III, Vincil, et al.
Pubblicazione: (2024)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
di: Zolduoarrati, Elijah, et al.
Pubblicazione: (2025)
di: Zolduoarrati, Elijah, et al.
Pubblicazione: (2025)
A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation
di: Qin, Qiaolin, et al.
Pubblicazione: (2025)
di: Qin, Qiaolin, et al.
Pubblicazione: (2025)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
di: Petrovic, Nenad, et al.
Pubblicazione: (2024)
di: Petrovic, Nenad, et al.
Pubblicazione: (2024)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
di: Sharma, Krishna Kumaar
Pubblicazione: (2025)
di: Sharma, Krishna Kumaar
Pubblicazione: (2025)
Large Language Models as Software Components: A Taxonomy for LLM-Integrated Applications
di: Weber, Irene
Pubblicazione: (2024)
di: Weber, Irene
Pubblicazione: (2024)
Learning Software Bug Reports: A Systematic Literature Review
di: Long, Guoming, et al.
Pubblicazione: (2025)
di: Long, Guoming, et al.
Pubblicazione: (2025)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
di: Eisenreich, Tobias, et al.
Pubblicazione: (2025)
di: Eisenreich, Tobias, et al.
Pubblicazione: (2025)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
Automated Code Review Using Large Language Models at Ericsson: An Experience Report
di: Ramesh, Shweta, et al.
Pubblicazione: (2025)
di: Ramesh, Shweta, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OODEval: Evaluating Large Language Models on Object-Oriented Design
di: Xiao, Bingxu, et al.
Pubblicazione: (2026) -
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
di: Alshaikh, Moaath, et al.
Pubblicazione: (2026) -
Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
di: Shah, Syed Muhammad Ashhar, et al.
Pubblicazione: (2026) -
How Quickly Do Development Teams Update Their Vulnerable Dependencies?
di: Rahman, Imranur, et al.
Pubblicazione: (2024) -
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
di: Paiva, Camila A., et al.
Pubblicazione: (2026)