Saved in:
| Main Authors: | Peters, Gideon, Khatoonabadi, SayedHassan, Shihab, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.05502 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Automated File-Level Logging Generation for Machine Learning Applications using LLMs: A Case Study using GPT-4o Mini
by: Rodriguez, Mayra Sofia Ruiz, et al.
Published: (2025)
by: Rodriguez, Mayra Sofia Ruiz, et al.
Published: (2025)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
by: Abedu, Samuel, et al.
Published: (2024)
by: Abedu, Samuel, et al.
Published: (2024)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)
by: Latendresse, Jasmine, et al.
Published: (2025)
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Is ChatGPT a Good Software Librarian? An Exploratory Study on the Use of ChatGPT for Software Library Recommendations
by: Latendresse, Jasmine, et al.
Published: (2024)
by: Latendresse, Jasmine, et al.
Published: (2024)
Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
by: Salim, Mohamad, et al.
Published: (2026)
by: Salim, Mohamad, et al.
Published: (2026)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
An Approach for Auto Generation of Labeling Functions for Software Engineering Chatbots
by: Alor, Ebube, et al.
Published: (2024)
by: Alor, Ebube, et al.
Published: (2024)
On Wasted Contributions: Understanding the Dynamics of Contributor-Abandoned Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
Predicting the First Response Latency of Maintainers and Contributors in Pull Requests
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
Understanding the Helpfulness of Stale Bot for Pull-based Development: An Empirical Study of 20 Large Open-Source Projects
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
by: Khatoonabadi, SayedHassan, et al.
Published: (2023)
The Impact of Environment Configurations on the Stability of AI-Enabled Systems
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
The Impact of Large Language Models (LLMs) on Code Review Process
by: Collante, Antonio, et al.
Published: (2025)
by: Collante, Antonio, et al.
Published: (2025)
Will It Survive? Deciphering the Fate of AI-Generated Code in Open Source
by: Rahman, Musfiqur, et al.
Published: (2026)
by: Rahman, Musfiqur, et al.
Published: (2026)
GAP2WSS: A Genetic Algorithm based on the Pareto Principle for Web Service Selection
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
by: Khatoonabadi, SayedHassan, et al.
Published: (2021)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Evaluating the Generalizability of LLMs in Automated Program Repair
by: Li, Fengjie, et al.
Published: (2025)
by: Li, Fengjie, et al.
Published: (2025)
The Present and Future of Bots in Software Engineering
by: Shihab, Emad, et al.
Published: (2022)
by: Shihab, Emad, et al.
Published: (2022)
Cybernaut: Towards Reliable Web Automation
by: Tomar, Ankur, et al.
Published: (2025)
by: Tomar, Ankur, et al.
Published: (2025)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
by: Bruches, Elena, et al.
Published: (2026)
by: Bruches, Elena, et al.
Published: (2026)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Harnessing the Power of LLMs: Automating Unit Test Generation for High-Performance Computing
by: Karanjai, Rabimba, et al.
Published: (2024)
by: Karanjai, Rabimba, et al.
Published: (2024)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
Utilizing LLMs for Industrial Process Automation
by: Fares, Salim
Published: (2026)
by: Fares, Salim
Published: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
by: Nunes, Henrique, et al.
Published: (2025)
by: Nunes, Henrique, et al.
Published: (2025)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
by: Lei, Xinping, et al.
Published: (2026)
by: Lei, Xinping, et al.
Published: (2026)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs
by: Hassan, Amna
Published: (2025)
by: Hassan, Amna
Published: (2025)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
FARM: Field-Aware Resolution Model for Intelligent Trigger-Action Automation
by: Badalov, Khusrav, et al.
Published: (2026)
by: Badalov, Khusrav, et al.
Published: (2026)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
WALL: A Web Application for Automated Quality Assurance using Large Language Models
by: Abtahi, Seyed Moein, et al.
Published: (2025)
by: Abtahi, Seyed Moein, et al.
Published: (2025)
Towards Automated Formal Verification of Backend Systems with LLMs
by: Xu, Kangping, et al.
Published: (2025)
by: Xu, Kangping, et al.
Published: (2025)
Similar Items
-
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025) -
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025) -
Automated File-Level Logging Generation for Machine Learning Applications using LLMs: A Case Study using GPT-4o Mini
by: Rodriguez, Mayra Sofia Ruiz, et al.
Published: (2025) -
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
by: Abedu, Samuel, et al.
Published: (2024) -
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)