SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
Fuente:
arXiv
Saved in:
| Main Authors: | Dinu, Ion George, Mihăescu, Marian Cristian, Rebedea, Traian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
by: Rao, Swanand
Published: (2026)
by: Rao, Swanand
Published: (2026)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
by: Rombaut, Benjamin
Published: (2026)
by: Rombaut, Benjamin
Published: (2026)
Proof of Concept as a First-Class Architectural Decision Instrument
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
by: Trooskens, Geert, et al.
Published: (2026)
by: Trooskens, Geert, et al.
Published: (2026)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
Technical Debt Management: The Road Ahead for Successful Software Delivery
by: Avgeriou, Paris, et al.
Published: (2024)
by: Avgeriou, Paris, et al.
Published: (2024)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
by: Yamasaki, Kazuma, et al.
Published: (2026)
by: Yamasaki, Kazuma, et al.
Published: (2026)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
by: Ranasinghe, Nishath Rajiv, et al.
Published: (2025)
by: Ranasinghe, Nishath Rajiv, et al.
Published: (2025)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
by: Costa, Igor
Published: (2026)
by: Costa, Igor
Published: (2026)
Characterizing JavaScript Security Code Smells
by: Kambhampati, Vikas, et al.
Published: (2024)
by: Kambhampati, Vikas, et al.
Published: (2024)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
An LLM-assisted approach to designing software architectures using ADD
by: Cervantes, Humberto, et al.
Published: (2025)
by: Cervantes, Humberto, et al.
Published: (2025)
Code Documentation and Analysis to Secure Software Development
by: Attie, Paul, et al.
Published: (2024)
by: Attie, Paul, et al.
Published: (2024)
xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation
by: Tran, Khoa A., et al.
Published: (2025)
by: Tran, Khoa A., et al.
Published: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
by: Calboreanu, Elias
Published: (2026)
by: Calboreanu, Elias
Published: (2026)
AgenticTyper: Automated Typing of Legacy Software Projects Using Agentic AI
by: Pohle, Clemens
Published: (2026)
by: Pohle, Clemens
Published: (2026)
Unified Modeling Language Code Generation from Diagram Images Using Multimodal Large Language Models
by: Bates, Averi, et al.
Published: (2025)
by: Bates, Averi, et al.
Published: (2025)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
by: Holt, Samuel, et al.
Published: (2023)
by: Holt, Samuel, et al.
Published: (2023)
AgentOps: Enabling Observability of LLM Agents
by: Dong, Liming, et al.
Published: (2024)
by: Dong, Liming, et al.
Published: (2024)
Helveg: Diagrams for Software Documentation
by: Štěpánek, Adam, et al.
Published: (2025)
by: Štěpánek, Adam, et al.
Published: (2025)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
From Scientific Texts to Verifiable Code: Automating the Process with Transformers
by: Wang, Changjie, et al.
Published: (2025)
by: Wang, Changjie, et al.
Published: (2025)
Quantitative Analysis of Technical Debt and Pattern Violation in Large Language Model Architectures
by: Slater, Tyler
Published: (2025)
by: Slater, Tyler
Published: (2025)
SPViz: A DSL-Driven Approach for Software Project Visualization Tooling
by: Rentz, Niklas, et al.
Published: (2024)
by: Rentz, Niklas, et al.
Published: (2024)
SAGAI-MID: A Generative AI-Driven Middleware for Dynamic Runtime Interoperability
by: Larsen, Oliver Aleksander, et al.
Published: (2026)
by: Larsen, Oliver Aleksander, et al.
Published: (2026)
Generative AI and the Transformation of Software Development Practices
by: Acharya, Vivek
Published: (2025)
by: Acharya, Vivek
Published: (2025)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026)
by: Boston, Marisa Ferrara, et al.
Published: (2026)
CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend
by: Meng, Haoming
Published: (2026)
by: Meng, Haoming
Published: (2026)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026)
by: Ashrafi, Nazmus
Published: (2026)
Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP
by: Vogel, Martin, et al.
Published: (2026)
by: Vogel, Martin, et al.
Published: (2026)
Similar Items
-
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
by: Rao, Swanand
Published: (2026) -
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
by: Rombaut, Benjamin
Published: (2026) -
Proof of Concept as a First-Class Architectural Decision Instrument
by: Antognolli, Bruno Fernando, et al.
Published: (2026) -
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025) -
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)