Creating benchmarkable components to measure the quality ofAI-enhanced developer tools
Fuente:
arXiv
Saved in:
| Main Authors: | Paradis, Elise, Murillo, Ambar, Pandey, Maulishree, D'Angelo, Sarah, Hughes, Matthew, Macvean, Andrew, Ferrari-Church, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How much does AI impact development speed? An enterprise-based randomized controlled trial
by: Paradis, Elise, et al.
Published: (2024)
by: Paradis, Elise, et al.
Published: (2024)
Foundational Requirements for Artificial General Intelligence: A Falsifiable Framework Based on Signal Prediction
by: Šprogar, Matej
Published: (2025)
by: Šprogar, Matej
Published: (2025)
OODEval: Evaluating Large Language Models on Object-Oriented Design
by: Xiao, Bingxu, et al.
Published: (2026)
by: Xiao, Bingxu, et al.
Published: (2026)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
by: Silva, Kaushitha, et al.
Published: (2026)
by: Silva, Kaushitha, et al.
Published: (2026)
ClustML: A Measure of Cluster Pattern Complexity in Scatterplots Learnt from Human-labeled Groupings
by: Abbas, Mostafa M., et al.
Published: (2021)
by: Abbas, Mostafa M., et al.
Published: (2021)
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
by: Pinchuk, Mykola
Published: (2026)
by: Pinchuk, Mykola
Published: (2026)
Scattered Forest Search: Smarter Code Space Exploration with LLMs
by: Light, Jonathan, et al.
Published: (2024)
by: Light, Jonathan, et al.
Published: (2024)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
BPMN to PDDL: Translating Business Workflows for AI Planning
by: Nie, Jasper, et al.
Published: (2025)
by: Nie, Jasper, et al.
Published: (2025)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
by: Palacios, Diego Cabezas
Published: (2026)
by: Palacios, Diego Cabezas
Published: (2026)
ScaleMAP: Preserving Local Density and Neighborhood Structure in Low-Dimensional Embeddings
by: Poorna, Rajas, et al.
Published: (2026)
by: Poorna, Rajas, et al.
Published: (2026)
Prior-Aligned Data Cleaning for Tabular Foundation Models
by: Berti-Equille, Laure
Published: (2026)
by: Berti-Equille, Laure
Published: (2026)
Using LLMs to Establish Implicit User Sentiment of Software Desirability
by: Weitl-Harms, Sherri, et al.
Published: (2024)
by: Weitl-Harms, Sherri, et al.
Published: (2024)
AI Model for Predicting Binding Affinity of Antidiabetic Compounds Targeting PPAR
by: Aman, La Ode, et al.
Published: (2024)
by: Aman, La Ode, et al.
Published: (2024)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)
by: Zolduoarrati, Elijah, et al.
Published: (2025)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
A Multidisciplinary Approach to Telegram Data Analysis
by: Varbanov, Velizar, et al.
Published: (2024)
by: Varbanov, Velizar, et al.
Published: (2024)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
Towards Generative Ray Path Sampling for Faster Point-to-Point Ray Tracing
by: Eertmans, Jérome, et al.
Published: (2024)
by: Eertmans, Jérome, et al.
Published: (2024)
Transform-Invariant Generative Ray Path Sampling for Efficient Radio Propagation Modeling
by: Eertmans, Jérome, et al.
Published: (2026)
by: Eertmans, Jérome, et al.
Published: (2026)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
by: Vicente-Sola, Alex, et al.
Published: (2022)
by: Vicente-Sola, Alex, et al.
Published: (2022)
Test Case Features as Hyper-heuristics for Inductive Programming
by: McDaid, Edward, et al.
Published: (2024)
by: McDaid, Edward, et al.
Published: (2024)
Reinforcement Learning for Dynamic Workflow Optimization in CI/CD Pipelines
by: Soni, Aniket Abhishek, et al.
Published: (2026)
by: Soni, Aniket Abhishek, et al.
Published: (2026)
Texterial: A Text-as-Material Interaction Paradigm for LLM-Mediated Writing
by: Shen, Jocelyn, et al.
Published: (2026)
by: Shen, Jocelyn, et al.
Published: (2026)
Model Generation with LLMs: From Requirements to UML Sequence Diagrams
by: Ferrari, Alessio, et al.
Published: (2024)
by: Ferrari, Alessio, et al.
Published: (2024)
Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale
by: Vaithilingam, Priyan, et al.
Published: (2025)
by: Vaithilingam, Priyan, et al.
Published: (2025)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026)
by: Ashrafi, Nazmus
Published: (2026)
Clover: A Neural-Symbolic Agentic Harness with Stochastic Tree-of-Thoughts for Verified RTL Repair
by: Luo, Zizhang, et al.
Published: (2026)
by: Luo, Zizhang, et al.
Published: (2026)
PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Instruction and Solution Probabilities as Heuristics for Inductive Programming
by: McDaid, Edward, et al.
Published: (2025)
by: McDaid, Edward, et al.
Published: (2025)
Transfer-Learning-Based Autotuning Using Gaussian Copula
by: Randall, Thomas, et al.
Published: (2024)
by: Randall, Thomas, et al.
Published: (2024)
A framework for realisable data-driven active flow control using model predictive control applied to a simplified truck wake
by: Solera-Rico, Alberto, et al.
Published: (2025)
by: Solera-Rico, Alberto, et al.
Published: (2025)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
by: Alshaikh, Moaath, et al.
Published: (2026)
by: Alshaikh, Moaath, et al.
Published: (2026)
A Tale of Two Systems: Characterizing Architectural Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
A Metrics-Oriented Architectural Model to Characterize Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
by: Kessel, Marcus
Published: (2025)
by: Kessel, Marcus
Published: (2025)
Exact Synthetic Populations for Scalable Societal and Market Modeling
by: Petit, Thierry, et al.
Published: (2025)
by: Petit, Thierry, et al.
Published: (2025)
Similar Items
-
How much does AI impact development speed? An enterprise-based randomized controlled trial
by: Paradis, Elise, et al.
Published: (2024) -
Foundational Requirements for Artificial General Intelligence: A Falsifiable Framework Based on Signal Prediction
by: Šprogar, Matej
Published: (2025) -
OODEval: Evaluating Large Language Models on Object-Oriented Design
by: Xiao, Bingxu, et al.
Published: (2026) -
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025) -
BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
by: Silva, Kaushitha, et al.
Published: (2026)