DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Xinhao, Yu, Jinke, Xu, Wenhao, Wen, Zeyi, Zhou, Ying, Liu, Junzhuo, Ji, Junhao, Chen, Zulong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911560589377536
author Huang, Xinhao
Yu, Jinke
Xu, Wenhao
Wen, Zeyi
Zhou, Ying
Liu, Junzhuo
Ji, Junhao
Chen, Zulong
author_facet Huang, Xinhao
Yu, Jinke
Xu, Wenhao
Wen, Zeyi
Zhou, Ying
Liu, Junzhuo
Ji, Junhao
Chen, Zulong
contents While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy with fine-grained visual details, often resulting in layout distortions or generic placeholders. To bridge this gap, we propose DOne, an end-to-end framework that decouples structure understanding from element rendering. DOne introduces (1) a learned layout segmentation module to decompose complex designs, avoiding the limitations of heuristic cropping; (2) a specialized hybrid element retriever to handle the extreme aspect ratios and densities of UI components; and (3) a schema-guided generation paradigm that bridges layout and code. To rigorously assess performance, we introduce HiFi2Code, a benchmark featuring significantly higher layout complexity than existing datasets. Extensive evaluations on the HiFi2Code demonstrate that DOne outperforms exiting methods in both high-level visual similarity (e.g., over 10% in GPT Score) and fine-grained element alignment. Human evaluations confirm a 3 times productivity gain with higher visual fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01226
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
Huang, Xinhao
Yu, Jinke
Xu, Wenhao
Wen, Zeyi
Zhou, Ying
Liu, Junzhuo
Ji, Junhao
Chen, Zulong
Computer Vision and Pattern Recognition
Software Engineering
While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy with fine-grained visual details, often resulting in layout distortions or generic placeholders. To bridge this gap, we propose DOne, an end-to-end framework that decouples structure understanding from element rendering. DOne introduces (1) a learned layout segmentation module to decompose complex designs, avoiding the limitations of heuristic cropping; (2) a specialized hybrid element retriever to handle the extreme aspect ratios and densities of UI components; and (3) a schema-guided generation paradigm that bridges layout and code. To rigorously assess performance, we introduce HiFi2Code, a benchmark featuring significantly higher layout complexity than existing datasets. Extensive evaluations on the HiFi2Code demonstrate that DOne outperforms exiting methods in both high-level visual similarity (e.g., over 10% in GPT Score) and fine-grained element alignment. Human evaluations confirm a 3 times productivity gain with higher visual fidelity.
title DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
topic Computer Vision and Pattern Recognition
Software Engineering
url https://arxiv.org/abs/2604.01226