Generating Automatic Feedback on UI Mockups with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Peitong, Warner, Jeremy, Li, Yang, Hartmann, Bjoern
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911804439920640
author Duan, Peitong
Warner, Jeremy
Li, Yang
Hartmann, Bjoern
author_facet Duan, Peitong
Warner, Jeremy
Li, Yang
Hartmann, Bjoern
contents Feedback on user interface (UI) mockups is crucial in design. However, human feedback is not always readily available. We explore the potential of using large language models for automatic feedback. Specifically, we focus on applying GPT-4 to automate heuristic evaluation, which currently entails a human expert assessing a UI's compliance with a set of design guidelines. We implemented a Figma plugin that takes in a UI design and a set of written heuristics, and renders automatically-generated feedback as constructive suggestions. We assessed performance on 51 UIs using three sets of guidelines, compared GPT-4-generated design suggestions with those from human experts, and conducted a study with 12 expert designers to understand fit with existing practice. We found that GPT-4-based feedback is useful for catching subtle errors, improving text, and considering UI semantics, but feedback also decreased in utility over iterations. Participants described several uses for this plugin despite its imperfect suggestions.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13139
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Automatic Feedback on UI Mockups with Large Language Models
Duan, Peitong
Warner, Jeremy
Li, Yang
Hartmann, Bjoern
Human-Computer Interaction
Feedback on user interface (UI) mockups is crucial in design. However, human feedback is not always readily available. We explore the potential of using large language models for automatic feedback. Specifically, we focus on applying GPT-4 to automate heuristic evaluation, which currently entails a human expert assessing a UI's compliance with a set of design guidelines. We implemented a Figma plugin that takes in a UI design and a set of written heuristics, and renders automatically-generated feedback as constructive suggestions. We assessed performance on 51 UIs using three sets of guidelines, compared GPT-4-generated design suggestions with those from human experts, and conducted a study with 12 expert designers to understand fit with existing practice. We found that GPT-4-based feedback is useful for catching subtle errors, improving text, and considering UI semantics, but feedback also decreased in utility over iterations. Participants described several uses for this plugin despite its imperfect suggestions.
title Generating Automatic Feedback on UI Mockups with Large Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2403.13139