Assessing LLMs for Front-end Software Architecture Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guerra, L. P. Franciscatto, Ernst, N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917949722329088
author Guerra, L. P. Franciscatto
Ernst, N.
author_facet Guerra, L. P. Franciscatto
Ernst, N.
contents Large Language Models (LLMs) have demonstrated significant promise in automating software development tasks, yet their capabilities with respect to software design tasks remains largely unclear. This study investigates the capabilities of an LLM in understanding, reproducing, and generating structures within the complex VIPER architecture, a design pattern for iOS applications. We leverage Bloom's taxonomy to develop a comprehensive evaluation framework to assess the LLM's performance across different cognitive domains such as remembering, understanding, applying, analyzing, evaluating, and creating. Experimental results, using ChatGPT 4 Turbo 2024-04-09, reveal that the LLM excelled in higher-order tasks like evaluating and creating, but faced challenges with lower-order tasks requiring precise retrieval of architectural details. These findings highlight both the potential of LLMs to reduce development costs and the barriers to their effective application in real-world software design scenarios. This study proposes a benchmark format for assessing LLM capabilities in software architecture, aiming to contribute toward more robust and accessible AI-driven development tools.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19518
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing LLMs for Front-end Software Architecture Knowledge
Guerra, L. P. Franciscatto
Ernst, N.
Software Engineering
Artificial Intelligence
D.2; I.2
Large Language Models (LLMs) have demonstrated significant promise in automating software development tasks, yet their capabilities with respect to software design tasks remains largely unclear. This study investigates the capabilities of an LLM in understanding, reproducing, and generating structures within the complex VIPER architecture, a design pattern for iOS applications. We leverage Bloom's taxonomy to develop a comprehensive evaluation framework to assess the LLM's performance across different cognitive domains such as remembering, understanding, applying, analyzing, evaluating, and creating. Experimental results, using ChatGPT 4 Turbo 2024-04-09, reveal that the LLM excelled in higher-order tasks like evaluating and creating, but faced challenges with lower-order tasks requiring precise retrieval of architectural details. These findings highlight both the potential of LLMs to reduce development costs and the barriers to their effective application in real-world software design scenarios. This study proposes a benchmark format for assessing LLM capabilities in software architecture, aiming to contribute toward more robust and accessible AI-driven development tools.
title Assessing LLMs for Front-end Software Architecture Knowledge
topic Software Engineering
Artificial Intelligence
D.2; I.2
url https://arxiv.org/abs/2502.19518