Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Mughal, Ali Hassaan, Fatima, Noor, Bilal, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
by: Sharma, Krishna Kumaar
Published: (2025)
by: Sharma, Krishna Kumaar
Published: (2025)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
by: Abualazm, Raafat, et al.
Published: (2026)
by: Abualazm, Raafat, et al.
Published: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven Agents
by: Gong, Jing, et al.
Published: (2024)
by: Gong, Jing, et al.
Published: (2024)
Interoperability From Kieker to OpenTelemetry: Demonstrated as Export to ExplorViz
by: Reichelt, David Georg, et al.
Published: (2024)
by: Reichelt, David Georg, et al.
Published: (2024)
SPViz: A DSL-Driven Approach for Software Project Visualization Tooling
by: Rentz, Niklas, et al.
Published: (2024)
by: Rentz, Niklas, et al.
Published: (2024)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
Overhead Measurement Noise in Different Runtime Environments
by: Reichelt, David Georg, et al.
Published: (2024)
by: Reichelt, David Georg, et al.
Published: (2024)
Characterising Contributions that Coincide with Vulnerability Mitigation in NPM Libraries
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
Energy-Aware Decision Making in Software Stack Upgrades
by: Stocker, Mirko, et al.
Published: (2026)
by: Stocker, Mirko, et al.
Published: (2026)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
by: Yeverechyahu, Doron, et al.
Published: (2024)
by: Yeverechyahu, Doron, et al.
Published: (2024)
Validating API Design Requirements for Interoperability: A Static Analysis Approach Using OpenAPI
by: Sundberg, Edwin, et al.
Published: (2025)
by: Sundberg, Edwin, et al.
Published: (2025)
Towards Identifying Code Proficiency through the Analysis of Python Textbooks
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
by: Rojpaisarnkit, Ruksit, et al.
Published: (2024)
Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
by: Zhong, Suzhen, et al.
Published: (2025)
by: Zhong, Suzhen, et al.
Published: (2025)
The Upper Bound of Information Diffusion in Code Review
by: Dorner, Michael, et al.
Published: (2023)
by: Dorner, Michael, et al.
Published: (2023)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Agentic Refactoring: An Empirical Study of AI Coding Agents
by: Horikawa, Kosei, et al.
Published: (2025)
by: Horikawa, Kosei, et al.
Published: (2025)
Migrating Esope to Fortran 2008 using model transformations
by: Sow, Younoussa, et al.
Published: (2026)
by: Sow, Younoussa, et al.
Published: (2026)
Choosing the Right Git Workflow: A Comparative Analysis of Trunk-based vs. Branch-based Approaches
by: Lopes, Pedro, et al.
Published: (2025)
by: Lopes, Pedro, et al.
Published: (2025)
LLM-based vs. Search-based Merge Conflict Resolution: An Empirical Study of Competing Paradigms
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
When Code Smells Meet ML: On the Lifecycle of ML-specific Code Smells in ML-enabled Systems
by: Recupito, Gilberto, et al.
Published: (2024)
by: Recupito, Gilberto, et al.
Published: (2024)
A Preliminary Study on Self-Contained Libraries in the NPM Ecosystem
by: Jaisri, Pongchai, et al.
Published: (2024)
by: Jaisri, Pongchai, et al.
Published: (2024)
A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation
by: Qin, Qiaolin, et al.
Published: (2025)
by: Qin, Qiaolin, et al.
Published: (2025)
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
by: Minh, Dao Sy Duy, et al.
Published: (2026)
by: Minh, Dao Sy Duy, et al.
Published: (2026)
Providing Information About Implemented Algorithms Improves Program Comprehension: A Controlled Experiment
by: Neumüller, Denis, et al.
Published: (2025)
by: Neumüller, Denis, et al.
Published: (2025)
Recommending Variable Names for Extract Local Variable Refactorings
by: Wang, Taiming, et al.
Published: (2025)
by: Wang, Taiming, et al.
Published: (2025)
Diagnosing Refactoring Dangers
by: Brinksma, Wouter, et al.
Published: (2024)
by: Brinksma, Wouter, et al.
Published: (2024)
A History Equivalence Algorithm for Dynamic Process Migration
by: Bakshi, Gargi, et al.
Published: (2024)
by: Bakshi, Gargi, et al.
Published: (2024)
Early-Stage Requirements Transformation Approaches: A Systematic Review
by: Letsholo, Keletso J.
Published: (2024)
by: Letsholo, Keletso J.
Published: (2024)
Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows
by: Shah, Syed Muhammad Ashhar, et al.
Published: (2026)
by: Shah, Syed Muhammad Ashhar, et al.
Published: (2026)
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
by: Paiva, Camila A., et al.
Published: (2026)
by: Paiva, Camila A., et al.
Published: (2026)
Comprehensive Evaluation of Large Language Models on Software Engineering Tasks: A Multi-Task Benchmark
by: Gunawan, Go Frendi, et al.
Published: (2026)
by: Gunawan, Go Frendi, et al.
Published: (2026)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
The Kieker Observability Framework Version 2
by: Yang, Shinhyung, et al.
Published: (2025)
by: Yang, Shinhyung, et al.
Published: (2025)
Automatic Generation of Conversational Interfaces for Tabular Data Analysis
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
Similar Items
-
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026) -
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
by: Sharma, Krishna Kumaar
Published: (2025) -
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
by: Abualazm, Raafat, et al.
Published: (2026) -
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026) -
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)