WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Zimu, Yang, Yunqiao, Ren, Houxing, Hou, Haotian, Xiao, Han, Wang, Ke, Shi, Weikang, Zhou, Aojun, Zhan, Mingjie, Li, Hongsheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
por: Lu, Zimu, et al.
Publicado: (2025)
por: Lu, Zimu, et al.
Publicado: (2025)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
por: Shi, Weikang, et al.
Publicado: (2026)
por: Shi, Weikang, et al.
Publicado: (2026)
Alignment with Fill-In-the-Middle for Enhancing Code Generation
por: Ren, Houxing, et al.
Publicado: (2025)
por: Ren, Houxing, et al.
Publicado: (2025)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
por: Yang, Yunqiao, et al.
Publicado: (2025)
por: Yang, Yunqiao, et al.
Publicado: (2025)
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
por: Yang, Yunqiao, et al.
Publicado: (2026)
por: Yang, Yunqiao, et al.
Publicado: (2026)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
por: Lu, Zimu, et al.
Publicado: (2024)
por: Lu, Zimu, et al.
Publicado: (2024)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
por: Ren, Houxing, et al.
Publicado: (2026)
por: Ren, Houxing, et al.
Publicado: (2026)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
por: Wang, Ke, et al.
Publicado: (2025)
por: Wang, Ke, et al.
Publicado: (2025)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
por: Lu, Zimu, et al.
Publicado: (2026)
por: Lu, Zimu, et al.
Publicado: (2026)
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
por: Lu, Zimu, et al.
Publicado: (2024)
por: Lu, Zimu, et al.
Publicado: (2024)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
por: Lu, Zimu, et al.
Publicado: (2024)
por: Lu, Zimu, et al.
Publicado: (2024)
Edit-Based Refinement for Parallel Masked Diffusion Language Models
por: Ren, Houxing, et al.
Publicado: (2026)
por: Ren, Houxing, et al.
Publicado: (2026)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
por: Wang, Ke, et al.
Publicado: (2025)
por: Wang, Ke, et al.
Publicado: (2025)
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
por: Ren, Houxing, et al.
Publicado: (2024)
por: Ren, Houxing, et al.
Publicado: (2024)
WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning
por: Jiang, Juyong, et al.
Publicado: (2026)
por: Jiang, Juyong, et al.
Publicado: (2026)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
por: Ren, Houxing, et al.
Publicado: (2024)
por: Ren, Houxing, et al.
Publicado: (2024)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
por: Wang, Ke, et al.
Publicado: (2024)
por: Wang, Ke, et al.
Publicado: (2024)
WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation
por: Wang, Kuang-Da, et al.
Publicado: (2025)
por: Wang, Kuang-Da, et al.
Publicado: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
por: Shi, Weikang, et al.
Publicado: (2025)
por: Shi, Weikang, et al.
Publicado: (2025)
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
por: Dai, Dasen, et al.
Publicado: (2026)
por: Dai, Dasen, et al.
Publicado: (2026)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
por: Hu, Yuxuan, et al.
Publicado: (2025)
por: Hu, Yuxuan, et al.
Publicado: (2025)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
por: Wang, Qiyao, et al.
Publicado: (2026)
por: Wang, Qiyao, et al.
Publicado: (2026)
NODI: Out-Of-Distribution Detection with Noise from Diffusion
por: Zhou, Jingqiu, et al.
Publicado: (2024)
por: Zhou, Jingqiu, et al.
Publicado: (2024)
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
por: Si, Yuan, et al.
Publicado: (2026)
por: Si, Yuan, et al.
Publicado: (2026)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
por: Ma, Liqun, et al.
Publicado: (2024)
por: Ma, Liqun, et al.
Publicado: (2024)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
por: Ramesh, Guruprasad Viswanathan, et al.
Publicado: (2026)
por: Ramesh, Guruprasad Viswanathan, et al.
Publicado: (2026)
RepoZero: Can LLMs Generate a Code Repository from Scratch?
por: Zhang, Zhaoxi, et al.
Publicado: (2026)
por: Zhang, Zhaoxi, et al.
Publicado: (2026)
ProgramBench: Can Language Models Rebuild Programs From Scratch?
por: Yang, John, et al.
Publicado: (2026)
por: Yang, John, et al.
Publicado: (2026)
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
por: Zhang, Hao, et al.
Publicado: (2026)
por: Zhang, Hao, et al.
Publicado: (2026)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
por: Guo, Zichun, et al.
Publicado: (2026)
por: Guo, Zichun, et al.
Publicado: (2026)
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
por: Yang, Haoyue, et al.
Publicado: (2026)
por: Yang, Haoyue, et al.
Publicado: (2026)
MolViBench: Evaluating LLMs on Molecular Vibe Coding
por: Li, Jiatong, et al.
Publicado: (2026)
por: Li, Jiatong, et al.
Publicado: (2026)
DriveSceneGen: Generating Diverse and Realistic Driving Scenarios from Scratch
por: Sun, Shuo, et al.
Publicado: (2023)
por: Sun, Shuo, et al.
Publicado: (2023)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
por: Lù, Xing Han, et al.
Publicado: (2024)
por: Lù, Xing Han, et al.
Publicado: (2024)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
por: Singh, Harman, et al.
Publicado: (2024)
por: Singh, Harman, et al.
Publicado: (2024)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
por: Li, Hanyu, et al.
Publicado: (2025)
por: Li, Hanyu, et al.
Publicado: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
por: Wang, Peng, et al.
Publicado: (2025)
por: Wang, Peng, et al.
Publicado: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
por: Xiao, Han, et al.
Publicado: (2025)
por: Xiao, Han, et al.
Publicado: (2025)
Measuring Web Accessibility Dimensions: An Evaluation of Understandability and Robustness in IIT Library Websites
por: Panda, Subhajit, et al.
Publicado: (2026)
por: Panda, Subhajit, et al.
Publicado: (2026)
Evian: Towards Explainable Visual Instruction-tuning Data Auditing
por: Jia, Zimu, et al.
Publicado: (2026)
por: Jia, Zimu, et al.
Publicado: (2026)
Ejemplares similares
-
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
por: Lu, Zimu, et al.
Publicado: (2025) -
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
por: Shi, Weikang, et al.
Publicado: (2026) -
Alignment with Fill-In-the-Middle for Enhancing Code Generation
por: Ren, Houxing, et al.
Publicado: (2025) -
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
por: Yang, Yunqiao, et al.
Publicado: (2025) -
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
por: Yang, Yunqiao, et al.
Publicado: (2026)