WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Shanchao, Jiang, Nan, Qian, Shangshu, Tan, Lin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908860620472320
author Liang, Shanchao
Jiang, Nan
Qian, Shangshu
Tan, Lin
author_facet Liang, Shanchao
Jiang, Nan
Qian, Shangshu
Tan, Lin
contents Web development involves turning UI designs into functional webpages, which can be difficult for both beginners and experienced developers due to the complexity of HTML's hierarchical structures and styles. While Large Language Models (LLMs) have shown promise in generating source code, two major challenges persist in UI-to-HTML code generation: (1) effectively representing HTML's hierarchical structure for LLMs, and (2) bridging the gap between the visual nature of UI designs and the text-based format of HTML code. To tackle these challenges, we introduce Waffle, a new fine-tuning strategy that uses a structure-aware attention mechanism to improve LLMs' understanding of HTML's structure and a contrastive fine-tuning approach to align LLMs' understanding of UI images and HTML code. Models fine-tuned with Waffle show up to 9.00 pp (percentage point) higher HTML match, 0.0982 higher CW-SSIM, 32.99 higher CLIP, and 27.12 pp higher LLEM on our new benchmark WebSight-Test and an existing benchmark Design2Code, outperforming current fine-tuning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18362
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
Liang, Shanchao
Jiang, Nan
Qian, Shangshu
Tan, Lin
Software Engineering
Computation and Language
Computer Vision and Pattern Recognition
Web development involves turning UI designs into functional webpages, which can be difficult for both beginners and experienced developers due to the complexity of HTML's hierarchical structures and styles. While Large Language Models (LLMs) have shown promise in generating source code, two major challenges persist in UI-to-HTML code generation: (1) effectively representing HTML's hierarchical structure for LLMs, and (2) bridging the gap between the visual nature of UI designs and the text-based format of HTML code. To tackle these challenges, we introduce Waffle, a new fine-tuning strategy that uses a structure-aware attention mechanism to improve LLMs' understanding of HTML's structure and a contrastive fine-tuning approach to align LLMs' understanding of UI images and HTML code. Models fine-tuned with Waffle show up to 9.00 pp (percentage point) higher HTML match, 0.0982 higher CW-SSIM, 32.99 higher CLIP, and 27.12 pp higher LLEM on our new benchmark WebSight-Test and an existing benchmark Design2Code, outperforming current fine-tuning methods.
title WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
topic Software Engineering
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.18362