UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jason, Schoop, Eldon, Leung, Alan, Barik, Titus, Bigham, Jeffrey P., Nichols, Jeffrey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909222156894208
author Wu, Jason
Schoop, Eldon
Leung, Alan
Barik, Titus
Bigham, Jeffrey P.
Nichols, Jeffrey
author_facet Wu, Jason
Schoop, Eldon
Leung, Alan
Barik, Titus
Bigham, Jeffrey P.
Nichols, Jeffrey
contents Large language models (LLMs) struggle to consistently generate UI code that compiles and produces visually relevant designs. Existing approaches to improve generation rely on expensive human feedback or distilling a proprietary model. In this paper, we explore the use of automated feedback (compilers and multi-modal models) to guide LLMs to generate high-quality UI code. Our method starts with an existing LLM and iteratively produces improved models by self-generating a large synthetic dataset using an original model, applying automated tools to aggressively filter, score, and de-duplicate the data into a refined higher quality dataset. The original LLM is improved by finetuning on this refined dataset. We applied our approach to several open-source LLMs and compared the resulting performance to baseline models with both automated metrics and human preferences. Our evaluation shows the resulting models outperform all other downloadable baselines and approach the performance of larger proprietary models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
Wu, Jason
Schoop, Eldon
Leung, Alan
Barik, Titus
Bigham, Jeffrey P.
Nichols, Jeffrey
Computation and Language
Human-Computer Interaction
Software Engineering
Large language models (LLMs) struggle to consistently generate UI code that compiles and produces visually relevant designs. Existing approaches to improve generation rely on expensive human feedback or distilling a proprietary model. In this paper, we explore the use of automated feedback (compilers and multi-modal models) to guide LLMs to generate high-quality UI code. Our method starts with an existing LLM and iteratively produces improved models by self-generating a large synthetic dataset using an original model, applying automated tools to aggressively filter, score, and de-duplicate the data into a refined higher quality dataset. The original LLM is improved by finetuning on this refined dataset. We applied our approach to several open-source LLMs and compared the resulting performance to baseline models with both automated metrics and human preferences. Our evaluation shows the resulting models outperform all other downloadable baselines and approach the performance of larger proprietary models.
title UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
topic Computation and Language
Human-Computer Interaction
Software Engineering
url https://arxiv.org/abs/2406.07739