WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Chenxu, Fu, Yingjie, Yang, Wei, Zhang, Ying, Xie, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
di: He, Yibo, et al.
Pubblicazione: (2025)
di: He, Yibo, et al.
Pubblicazione: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
di: Cui, Yi
Pubblicazione: (2024)
di: Cui, Yi
Pubblicazione: (2024)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
di: Lei, Xinping, et al.
Pubblicazione: (2026)
di: Lei, Xinping, et al.
Pubblicazione: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
di: Gao, Zeyu, et al.
Pubblicazione: (2025)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
di: Bai, Haoyue, et al.
Pubblicazione: (2026)
di: Bai, Haoyue, et al.
Pubblicazione: (2026)
WebSuite: Systematically Evaluating Why Web Agents Fail
di: Li, Eric, et al.
Pubblicazione: (2024)
di: Li, Eric, et al.
Pubblicazione: (2024)
PerfCoder: Large Language Models for Interpretable Code Performance Optimization
di: Yang, Jiuding, et al.
Pubblicazione: (2025)
di: Yang, Jiuding, et al.
Pubblicazione: (2025)
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
di: Yin, Wenjing, et al.
Pubblicazione: (2025)
di: Yin, Wenjing, et al.
Pubblicazione: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
di: Kong, Fanheng, et al.
Pubblicazione: (2026)
di: Kong, Fanheng, et al.
Pubblicazione: (2026)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
di: Xiao, Jingyu, et al.
Pubblicazione: (2025)
di: Xiao, Jingyu, et al.
Pubblicazione: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
di: Li, Chunyang, et al.
Pubblicazione: (2025)
di: Li, Chunyang, et al.
Pubblicazione: (2025)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
di: Cui, Yi
Pubblicazione: (2024)
di: Cui, Yi
Pubblicazione: (2024)
Investigating Training Data Detection in AI Coders
di: Li, Tianlin, et al.
Pubblicazione: (2025)
di: Li, Tianlin, et al.
Pubblicazione: (2025)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
di: He, Zehai, et al.
Pubblicazione: (2026)
di: He, Zehai, et al.
Pubblicazione: (2026)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
di: Xu, Mingde, et al.
Pubblicazione: (2025)
di: Xu, Mingde, et al.
Pubblicazione: (2025)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2024)
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
di: Liu, Chenxu, et al.
Pubblicazione: (2025)
di: Liu, Chenxu, et al.
Pubblicazione: (2025)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
di: Wu, Yifan, et al.
Pubblicazione: (2026)
di: Wu, Yifan, et al.
Pubblicazione: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
di: Tóth, Rebeka, et al.
Pubblicazione: (2024)
di: Tóth, Rebeka, et al.
Pubblicazione: (2024)
WybeCoder: Verified Imperative Code Generation
di: Gloeckle, Fabian, et al.
Pubblicazione: (2026)
di: Gloeckle, Fabian, et al.
Pubblicazione: (2026)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
di: Wan, Yuxuan, et al.
Pubblicazione: (2025)
di: Wan, Yuxuan, et al.
Pubblicazione: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
di: Yang, Ruozhao, et al.
Pubblicazione: (2026)
di: Yang, Ruozhao, et al.
Pubblicazione: (2026)
StarCoder 2 and The Stack v2: The Next Generation
di: Lozhkov, Anton, et al.
Pubblicazione: (2024)
di: Lozhkov, Anton, et al.
Pubblicazione: (2024)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
di: Ma, Chenyang, et al.
Pubblicazione: (2026)
di: Ma, Chenyang, et al.
Pubblicazione: (2026)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
di: Tran, Hung, et al.
Pubblicazione: (2026)
di: Tran, Hung, et al.
Pubblicazione: (2026)
From PowerPoint UI Sketches to Web-Based Applications: Pattern-Driven Code Generation for GIS Dashboard Development Using Knowledge-Augmented LLMs, Context-Aware Visual Prompting, and the React Framework
di: Xu, Haowen, et al.
Pubblicazione: (2025)
di: Xu, Haowen, et al.
Pubblicazione: (2025)
FastCoder: Accelerating Repository-level Code Generation via Efficient Retrieval and Verification
di: Zhao, Qianhui, et al.
Pubblicazione: (2025)
di: Zhao, Qianhui, et al.
Pubblicazione: (2025)
o1-Coder: an o1 Replication for Coding
di: Zhang, Yuxiang, et al.
Pubblicazione: (2024)
di: Zhang, Yuxiang, et al.
Pubblicazione: (2024)
MemoCoder: Automated Function Synthesis using LLM-Supported Agents
di: Jia, Yiping, et al.
Pubblicazione: (2025)
di: Jia, Yiping, et al.
Pubblicazione: (2025)
InCoder-32B: Code Foundation Model for Industrial Scenarios
di: Yang, Jian, et al.
Pubblicazione: (2026)
di: Yang, Jian, et al.
Pubblicazione: (2026)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
di: Peters, Gideon, et al.
Pubblicazione: (2026)
di: Peters, Gideon, et al.
Pubblicazione: (2026)
WALL: A Web Application for Automated Quality Assurance using Large Language Models
di: Abtahi, Seyed Moein, et al.
Pubblicazione: (2025)
di: Abtahi, Seyed Moein, et al.
Pubblicazione: (2025)
BabelCoder: Agentic Code Translation with Specification Alignment
di: Rabbi, Fazle, et al.
Pubblicazione: (2025)
di: Rabbi, Fazle, et al.
Pubblicazione: (2025)
Cybernaut: Towards Reliable Web Automation
di: Tomar, Ankur, et al.
Pubblicazione: (2025)
di: Tomar, Ankur, et al.
Pubblicazione: (2025)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
di: Garg, Spandan, et al.
Pubblicazione: (2025)
di: Garg, Spandan, et al.
Pubblicazione: (2025)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
di: He, Yibo, et al.
Pubblicazione: (2025) -
FullStack Bench: Evaluating LLMs as Full Stack Coders
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024) -
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
di: Cui, Yi
Pubblicazione: (2024) -
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
di: Lei, Xinping, et al.
Pubblicazione: (2026) -
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
di: Gao, Zeyu, et al.
Pubblicazione: (2025)