DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Jingyu, Wang, Ming, Lam, Man Ho, Wan, Yuxuan, Liu, Junliang, Huo, Yintong, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
by: Xiao, Jingyu, et al.
Published: (2026)
by: Xiao, Jingyu, et al.
Published: (2026)
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
by: Xiao, Jingyu, et al.
Published: (2024)
by: Xiao, Jingyu, et al.
Published: (2024)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements
by: Wan, Yuxuan, et al.
Published: (2026)
by: Wan, Yuxuan, et al.
Published: (2026)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
by: Wan, Yuxuan, et al.
Published: (2025)
by: Wan, Yuxuan, et al.
Published: (2025)
UIBenchKit: A unified toolkit for design-to-code model evaluation
by: Le, Chinh T., et al.
Published: (2026)
by: Le, Chinh T., et al.
Published: (2026)
Envisioning Future Interactive Web Development: Editing Webpage with Natural Language
by: Dang, Truong Hai, et al.
Published: (2025)
by: Dang, Truong Hai, et al.
Published: (2025)
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Prototype2Code: End-to-end Front-end Code Generation from UI Design Prototypes
by: Xiao, Shuhong, et al.
Published: (2024)
by: Xiao, Shuhong, et al.
Published: (2024)
90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development
by: Yang, Runxin, et al.
Published: (2025)
by: Yang, Runxin, et al.
Published: (2025)
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Single-Language Evidence Is Insufficient for Automated Logging: A Multilingual Benchmark and Empirical Study with LLMs
by: Zhong, Renyi, et al.
Published: (2026)
by: Zhong, Renyi, et al.
Published: (2026)
Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
by: Lu, Ruofan, et al.
Published: (2025)
by: Lu, Ruofan, et al.
Published: (2025)
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
by: Wang, Minxing, et al.
Published: (2026)
by: Wang, Minxing, et al.
Published: (2026)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026)
by: Lam, Man Ho, et al.
Published: (2026)
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
by: Lam, Man Ho, et al.
Published: (2025)
by: Lam, Man Ho, et al.
Published: (2025)
End-to-End Automated Logging via Multi-Agent Framework
by: Zhong, Renyi, et al.
Published: (2025)
by: Zhong, Renyi, et al.
Published: (2025)
Small is Beautiful: A Practical and Efficient Log Parsing Framework
by: Wang, Minxing, et al.
Published: (2026)
by: Wang, Minxing, et al.
Published: (2026)
CodeAD: Synthesize Code of Rules for Log-based Anomaly Detection with LLMs
by: Huang, Junjie, et al.
Published: (2025)
by: Huang, Junjie, et al.
Published: (2025)
Larger Is Not Always Better: Exploring Small Open-source Language Models in Logging Statement Generation
by: Zhong, Renyi, et al.
Published: (2025)
by: Zhong, Renyi, et al.
Published: (2025)
CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency
by: Zhong, Renyi, et al.
Published: (2025)
by: Zhong, Renyi, et al.
Published: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
Modular Layout Synthesis (MLS): Front-end Code via Structure Normalization and Constrained Generation
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
LogUpdater: Automated Detection and Repair of Specific Defects in Logging Statements
by: Zhong, Renyi, et al.
Published: (2024)
by: Zhong, Renyi, et al.
Published: (2024)
Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study
by: Li, Yichen, et al.
Published: (2023)
by: Li, Yichen, et al.
Published: (2023)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
by: Lu, Ruofan, et al.
Published: (2025)
by: Lu, Ruofan, et al.
Published: (2025)
AnomalyGen: Enhancing Log-Based Anomaly Detection with Code-Guided Data Augmentation
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
COFFE: A Code Efficiency Benchmark for Code Generation
by: Peng, Yun, et al.
Published: (2025)
by: Peng, Yun, et al.
Published: (2025)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
by: Guo, Xiaoyu, et al.
Published: (2025)
by: Guo, Xiaoyu, et al.
Published: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
by: Ouyang, Shuyin, et al.
Published: (2025)
by: Ouyang, Shuyin, et al.
Published: (2025)
Go Static: Contextualized Logging Statement Generation
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
Prompting for Automatic Log Template Extraction
by: Xu, Junjielong, et al.
Published: (2023)
by: Xu, Junjielong, et al.
Published: (2023)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
by: Jeon, YoungHoon, et al.
Published: (2026)
by: Jeon, YoungHoon, et al.
Published: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Similar Items
-
EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
by: Xiao, Jingyu, et al.
Published: (2025) -
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
by: Xiao, Jingyu, et al.
Published: (2026) -
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
by: Xiao, Jingyu, et al.
Published: (2024) -
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024) -
From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements
by: Wan, Yuxuan, et al.
Published: (2026)