DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911560589377536 |
|---|---|
| author | Huang, Xinhao Yu, Jinke Xu, Wenhao Wen, Zeyi Zhou, Ying Liu, Junzhuo Ji, Junhao Chen, Zulong |
| author_facet | Huang, Xinhao Yu, Jinke Xu, Wenhao Wen, Zeyi Zhou, Ying Liu, Junzhuo Ji, Junhao Chen, Zulong |
| contents | While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy with fine-grained visual details, often resulting in layout distortions or generic placeholders. To bridge this gap, we propose DOne, an end-to-end framework that decouples structure understanding from element rendering. DOne introduces (1) a learned layout segmentation module to decompose complex designs, avoiding the limitations of heuristic cropping; (2) a specialized hybrid element retriever to handle the extreme aspect ratios and densities of UI components; and (3) a schema-guided generation paradigm that bridges layout and code. To rigorously assess performance, we introduce HiFi2Code, a benchmark featuring significantly higher layout complexity than existing datasets. Extensive evaluations on the HiFi2Code demonstrate that DOne outperforms exiting methods in both high-level visual similarity (e.g., over 10% in GPT Score) and fine-grained element alignment. Human evaluations confirm a 3 times productivity gain with higher visual fidelity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_01226 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation Huang, Xinhao Yu, Jinke Xu, Wenhao Wen, Zeyi Zhou, Ying Liu, Junzhuo Ji, Junhao Chen, Zulong Computer Vision and Pattern Recognition Software Engineering While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy with fine-grained visual details, often resulting in layout distortions or generic placeholders. To bridge this gap, we propose DOne, an end-to-end framework that decouples structure understanding from element rendering. DOne introduces (1) a learned layout segmentation module to decompose complex designs, avoiding the limitations of heuristic cropping; (2) a specialized hybrid element retriever to handle the extreme aspect ratios and densities of UI components; and (3) a schema-guided generation paradigm that bridges layout and code. To rigorously assess performance, we introduce HiFi2Code, a benchmark featuring significantly higher layout complexity than existing datasets. Extensive evaluations on the HiFi2Code demonstrate that DOne outperforms exiting methods in both high-level visual similarity (e.g., over 10% in GPT Score) and fine-grained element alignment. Human evaluations confirm a 3 times productivity gain with higher visual fidelity. |
| title | DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation |
| topic | Computer Vision and Pattern Recognition Software Engineering |
| url | https://arxiv.org/abs/2604.01226 |