Toward Architecture-Aware Evaluation Metrics for LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Souza, Débora, Machado, Patrícia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
von: He, Kaifeng, et al.
Veröffentlicht: (2025)
von: He, Kaifeng, et al.
Veröffentlicht: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
von: Liu, Shunyu, et al.
Veröffentlicht: (2025)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
von: Sellami, Khaled, et al.
Veröffentlicht: (2025)
von: Sellami, Khaled, et al.
Veröffentlicht: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026)
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026)
OODEval: Evaluating Large Language Models on Object-Oriented Design
von: Xiao, Bingxu, et al.
Veröffentlicht: (2026)
von: Xiao, Bingxu, et al.
Veröffentlicht: (2026)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
von: Kessel, Marcus
Veröffentlicht: (2025)
von: Kessel, Marcus
Veröffentlicht: (2025)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
von: Trooskens, Geert, et al.
Veröffentlicht: (2026)
von: Trooskens, Geert, et al.
Veröffentlicht: (2026)
CIFE: Code Instruction-Following Evaluation
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
von: Gunnu, Sravani, et al.
Veröffentlicht: (2025)
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025)
von: Long, Guoming, et al.
Veröffentlicht: (2025)
Automating Domain-Driven Design: Experience with a Prompting Framework
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2026)
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2025)
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2025)
Can AI Assist in Olympiad Coding
von: Ren, Samuel
Veröffentlicht: (2025)
von: Ren, Samuel
Veröffentlicht: (2025)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
von: Abualazm, Raafat, et al.
Veröffentlicht: (2026)
von: Abualazm, Raafat, et al.
Veröffentlicht: (2026)
EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
von: Costa, Igor, et al.
Veröffentlicht: (2025)
von: Costa, Igor, et al.
Veröffentlicht: (2025)
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
von: Alam, Khairul, et al.
Veröffentlicht: (2024)
von: Alam, Khairul, et al.
Veröffentlicht: (2024)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
von: Bradbury, Jeremy S., et al.
Veröffentlicht: (2024)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
von: Kessel, Marcus
Veröffentlicht: (2024)
von: Kessel, Marcus
Veröffentlicht: (2024)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
von: Sharma, Krishna Kumaar
Veröffentlicht: (2025)
von: Sharma, Krishna Kumaar
Veröffentlicht: (2025)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
von: Lebioda, Krzysztof, et al.
Veröffentlicht: (2024)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
von: Ranasinghe, Nishath Rajiv, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Nishath Rajiv, et al.
Veröffentlicht: (2025)
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures
von: Rombaut, Benjamin
Veröffentlicht: (2026)
von: Rombaut, Benjamin
Veröffentlicht: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
von: Dinu, Ion George, et al.
Veröffentlicht: (2026)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
von: Petrovic, Nenad, et al.
Veröffentlicht: (2024)
von: Petrovic, Nenad, et al.
Veröffentlicht: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
N-Version Assessment and Enhancement of Generative AI
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
Morescient GAI for Software Engineering (Extended Version)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
von: Kessel, Marcus, et al.
Veröffentlicht: (2024)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
Unified Modeling Language Code Generation from Diagram Images Using Multimodal Large Language Models
von: Bates, Averi, et al.
Veröffentlicht: (2025)
von: Bates, Averi, et al.
Veröffentlicht: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
von: Mughal, Ali Hassaan, et al.
Veröffentlicht: (2026)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
Plan with Code: Comparing approaches for robust NL to DSL generation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
von: Young, Richard J.
Veröffentlicht: (2025)
von: Young, Richard J.
Veröffentlicht: (2025)
Circularity and Symmetries of $p$ and $p^{2}$-polygons
von: Haag, Rolf
Veröffentlicht: (2025)
von: Haag, Rolf
Veröffentlicht: (2025)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
von: Yeverechyahu, Doron, et al.
Veröffentlicht: (2024)
von: Yeverechyahu, Doron, et al.
Veröffentlicht: (2024)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests
von: Yamasaki, Kazuma, et al.
Veröffentlicht: (2026)
von: Yamasaki, Kazuma, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
von: He, Kaifeng, et al.
Veröffentlicht: (2025) -
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025) -
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
von: Liu, Shunyu, et al.
Veröffentlicht: (2025) -
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
von: Sellami, Khaled, et al.
Veröffentlicht: (2025) -
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
von: Ahmed, Sheikh Nazib, et al.
Veröffentlicht: (2026)