ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adnan, Bassam, Gupta, Aviral, Akshathala, Sreemaee, Vaidhyanathan, Karthik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
von: Akshathala, Sreemaee, et al.
Veröffentlicht: (2025)
von: Akshathala, Sreemaee, et al.
Veröffentlicht: (2025)
Can AI Agents Generate Microservices? How Far are We?
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
Context Matters: Evaluating Context Strategies for Automated ADR Generation Using LLMs
von: Gupta, Aviral, et al.
Veröffentlicht: (2026)
von: Gupta, Aviral, et al.
Veröffentlicht: (2026)
Leveraging LLMs for Dynamic IoT Systems Generation through Mixed-Initiative Interaction
von: Adnan, Bassam, et al.
Veröffentlicht: (2025)
von: Adnan, Bassam, et al.
Veröffentlicht: (2025)
LLMs for Generation of Architectural Components: An Exploratory Empirical Study in the Serverless World
von: Arun, Shrikara, et al.
Veröffentlicht: (2025)
von: Arun, Shrikara, et al.
Veröffentlicht: (2025)
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
von: Parthasarathy, Kannan, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Kannan, et al.
Veröffentlicht: (2025)
AgenticAKM : Enroute to Agentic Architecture Knowledge Management
von: Dhar, Rudra, et al.
Veröffentlicht: (2026)
von: Dhar, Rudra, et al.
Veröffentlicht: (2026)
LLM-based Automated Architecture View Generation: Where Are We Now?
von: Sathvika, Miryala, et al.
Veröffentlicht: (2026)
von: Sathvika, Miryala, et al.
Veröffentlicht: (2026)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
von: Dhar, Rudra, et al.
Veröffentlicht: (2024)
von: Dhar, Rudra, et al.
Veröffentlicht: (2024)
POLARIS: Is Multi-Agentic Reasoning the Next Wave in Engineering Self-Adaptive Systems?
von: Pandey, Divyansh, et al.
Veröffentlicht: (2025)
von: Pandey, Divyansh, et al.
Veröffentlicht: (2025)
Approach Towards Semi-Automated Certification for Low Criticality ML-Enabled Airborne Applications
von: Sridhar, Chandrasekar, et al.
Veröffentlicht: (2025)
von: Sridhar, Chandrasekar, et al.
Veröffentlicht: (2025)
Generative AI for Software Architecture. Applications, Challenges, and Future Directions
von: Esposito, Matteo, et al.
Veröffentlicht: (2025)
von: Esposito, Matteo, et al.
Veröffentlicht: (2025)
CALM: A Self-Adaptive Orchestration Approach for QoS-Aware Routing in Small Language Model based Systems
von: Jain, Hemang, et al.
Veröffentlicht: (2026)
von: Jain, Hemang, et al.
Veröffentlicht: (2026)
Architecting AgentOps Needs CHANGE
von: Biswas, Shaunak, et al.
Veröffentlicht: (2026)
von: Biswas, Shaunak, et al.
Veröffentlicht: (2026)
EnCoDe: Energy Estimation of Source Code At Design-Time
von: Goyal, Shailender, et al.
Veröffentlicht: (2026)
von: Goyal, Shailender, et al.
Veröffentlicht: (2026)
Toward architecting self-coding information systems
von: Falcão, Rodrigo, et al.
Veröffentlicht: (2026)
von: Falcão, Rodrigo, et al.
Veröffentlicht: (2026)
EcoMLS: A Self-Adaptation Approach for Architecting Green ML-Enabled Systems
von: Tedla, Meghana, et al.
Veröffentlicht: (2024)
von: Tedla, Meghana, et al.
Veröffentlicht: (2024)
SWITCH: An Exemplar for Evaluating Self-Adaptive ML-Enabled Systems
von: Marda, Arya, et al.
Veröffentlicht: (2024)
von: Marda, Arya, et al.
Veröffentlicht: (2024)
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
EdgeMLBalancer: A Self-Adaptive Approach for Dynamic Model Switching on Resource-Constrained Edge Devices
von: Matathammal, Akhila, et al.
Veröffentlicht: (2025)
von: Matathammal, Akhila, et al.
Veröffentlicht: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
ArchAgent: Scalable Legacy Software Architecture Recovery with LLMs
von: Pan, Rusheng, et al.
Veröffentlicht: (2026)
von: Pan, Rusheng, et al.
Veröffentlicht: (2026)
Towards Architecting Sustainable MLOps: A Self-Adaptation Approach
von: Bhatt, Hiya, et al.
Veröffentlicht: (2024)
von: Bhatt, Hiya, et al.
Veröffentlicht: (2024)
Reimagining Self-Adaptation in the Age of Large Language Models
von: Donakanti, Raghav, et al.
Veröffentlicht: (2024)
von: Donakanti, Raghav, et al.
Veröffentlicht: (2024)
Dissecting Transformers: A CLEAR Perspective towards Green AI
von: Jain, Hemang, et al.
Veröffentlicht: (2025)
von: Jain, Hemang, et al.
Veröffentlicht: (2025)
DRAFT-ing Architectural Design Decisions using LLMs
von: Dhar, Rudra, et al.
Veröffentlicht: (2025)
von: Dhar, Rudra, et al.
Veröffentlicht: (2025)
SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
von: Tripathy, Arihant, et al.
Veröffentlicht: (2025)
von: Tripathy, Arihant, et al.
Veröffentlicht: (2025)
APDT: A Digital Twin for Assessing Access Point Characteristics in a Network
von: Yashaswinee, D. Sree, et al.
Veröffentlicht: (2025)
von: Yashaswinee, D. Sree, et al.
Veröffentlicht: (2025)
DyPyBench: A Benchmark of Executable Python Software
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
HarmonE: A Self-Adaptive Approach to Architecting Sustainable MLOps
von: Bhatt, Hiya, et al.
Veröffentlicht: (2025)
von: Bhatt, Hiya, et al.
Veröffentlicht: (2025)
Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2026)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2026)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
LoCoML: A Framework for Real-World ML Inference Pipelines
von: Maddireddy, Kritin, et al.
Veröffentlicht: (2025)
von: Maddireddy, Kritin, et al.
Veröffentlicht: (2025)
Generating Energy-Efficient Code via Large-Language Models -- Where are we now?
von: Apsan, Radu, et al.
Veröffentlicht: (2025)
von: Apsan, Radu, et al.
Veröffentlicht: (2025)
Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
von: Das, Gunjan, et al.
Veröffentlicht: (2025)
von: Das, Gunjan, et al.
Veröffentlicht: (2025)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
TwinArch: A Digital Twin Reference Architecture
von: Somma, Alessandra, et al.
Veröffentlicht: (2025)
von: Somma, Alessandra, et al.
Veröffentlicht: (2025)
RubberDuckBench: A Benchmark for AI Coding Assistants
von: Mohammed, Ferida, et al.
Veröffentlicht: (2026)
von: Mohammed, Ferida, et al.
Veröffentlicht: (2026)
Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering
von: Bhati, Happy
Veröffentlicht: (2026)
von: Bhati, Happy
Veröffentlicht: (2026)
Harmonica: A Self-Adaptation Exemplar for Sustainable MLOps
von: Halgatti, Ananya, et al.
Veröffentlicht: (2026)
von: Halgatti, Ananya, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
von: Akshathala, Sreemaee, et al.
Veröffentlicht: (2025) -
Can AI Agents Generate Microservices? How Far are We?
von: Adnan, Bassam, et al.
Veröffentlicht: (2026) -
Context Matters: Evaluating Context Strategies for Automated ADR Generation Using LLMs
von: Gupta, Aviral, et al.
Veröffentlicht: (2026) -
Leveraging LLMs for Dynamic IoT Systems Generation through Mixed-Initiative Interaction
von: Adnan, Bassam, et al.
Veröffentlicht: (2025) -
LLMs for Generation of Architectural Components: An Exploratory Empirical Study in the Serverless World
von: Arun, Shrikara, et al.
Veröffentlicht: (2025)