Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Chenglei, Zhang, Yanzhe, Li, Ryan, Yang, Zhengyuan, Liu, Ruibo, Yang, Diyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929704350515200
author Si, Chenglei
Zhang, Yanzhe
Li, Ryan
Yang, Zhengyuan
Liu, Ruibo
Yang, Diyi
author_facet Si, Chenglei
Zhang, Yanzhe
Li, Ryan
Yang, Zhengyuan
Liu, Ruibo
Yang, Diyi
contents Generative AI has made rapid advancements in recent years, achieving unprecedented capabilities in multimodal understanding and code generation. This can enable a new paradigm of front-end development in which multimodal large language models (MLLMs) directly convert visual designs into code implementations. In this work, we construct Design2Code - the first real-world benchmark for this task. Specifically, we manually curate 484 diverse real-world webpages as test cases and develop a set of automatic evaluation metrics to assess how well current multimodal LLMs can generate the code implementations that directly render into the given reference webpages, given the screenshots as input. We also complement automatic metrics with comprehensive human evaluations to validate the performance ranking. To rigorously benchmark MLLMs, we test various multimodal prompting methods on frontier models such as GPT-4o, GPT-4V, Gemini, and Claude. Our fine-grained break-down metrics indicate that models mostly lag in recalling visual elements from the input webpages and generating correct layout designs.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03163
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Si, Chenglei
Zhang, Yanzhe
Li, Ryan
Yang, Zhengyuan
Liu, Ruibo
Yang, Diyi
Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
Generative AI has made rapid advancements in recent years, achieving unprecedented capabilities in multimodal understanding and code generation. This can enable a new paradigm of front-end development in which multimodal large language models (MLLMs) directly convert visual designs into code implementations. In this work, we construct Design2Code - the first real-world benchmark for this task. Specifically, we manually curate 484 diverse real-world webpages as test cases and develop a set of automatic evaluation metrics to assess how well current multimodal LLMs can generate the code implementations that directly render into the given reference webpages, given the screenshots as input. We also complement automatic metrics with comprehensive human evaluations to validate the performance ranking. To rigorously benchmark MLLMs, we test various multimodal prompting methods on frontier models such as GPT-4o, GPT-4V, Gemini, and Claude. Our fine-grained break-down metrics indicate that models mostly lag in recalling visual elements from the input webpages and generating correct layout designs.
title Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
topic Computation and Language
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2403.03163