Similar Items
Insights from Benchmarking Frontier Language Models on Web App Code Generation
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
Rapid Mobile App Development for Generative AI Agents on MIT App Inventor
by: Gao, Jaida, et al.
Published: (2024)
by: Gao, Jaida, et al.
Published: (2024)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)
by: Cui, Yi
Published: (2025)
WhatsCode: Large-Scale GenAI Deployment for Developer Efficiency at WhatsApp
by: Mao, Ke, et al.
Published: (2025)
by: Mao, Ke, et al.
Published: (2025)
AppForge: From Assistant to Independent Developer -- Are GPTs Ready for Software Development?
by: Ran, Dezhi, et al.
Published: (2025)
by: Ran, Dezhi, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
by: Lei, Xinping, et al.
Published: (2026)
by: Lei, Xinping, et al.
Published: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
by: Xu, Mingde, et al.
Published: (2025)
by: Xu, Mingde, et al.
Published: (2025)
Skeet: Towards a Lightweight Serverless Framework Supporting Modern AI-Driven App Development
by: Fumitake, Kawasaki, et al.
Published: (2024)
by: Fumitake, Kawasaki, et al.
Published: (2024)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026)
by: Bai, Haoyue, et al.
Published: (2026)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
by: Trivedi, Harsh, et al.
Published: (2024)
by: Trivedi, Harsh, et al.
Published: (2024)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
by: He, Zehai, et al.
Published: (2026)
by: He, Zehai, et al.
Published: (2026)
app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
by: Kniazev, Evgenii, et al.
Published: (2025)
by: Kniazev, Evgenii, et al.
Published: (2025)
A Dual-Helix Governance Approach Towards Reliable Agentic AI for WebGIS Development
by: Boyuan, et al.
Published: (2026)
by: Boyuan, et al.
Published: (2026)
From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems
by: Ortiz, Marcos, et al.
Published: (2025)
by: Ortiz, Marcos, et al.
Published: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Getting Inspiration for Feature Elicitation: App Store- vs. LLM-based Approach
by: Wei, Jialiang, et al.
Published: (2024)
by: Wei, Jialiang, et al.
Published: (2024)
Exploring Zero-Shot App Review Classification with ChatGPT: Challenges and Potential
by: Chaudhary, Mohit, et al.
Published: (2025)
by: Chaudhary, Mohit, et al.
Published: (2025)
AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
On Developers' Self-Declaration of AI-Generated Code: An Analysis of Practices
by: Kashif, Syed Mohammad, et al.
Published: (2025)
by: Kashif, Syed Mohammad, et al.
Published: (2025)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
by: Wan, Yuxuan, et al.
Published: (2025)
by: Wan, Yuxuan, et al.
Published: (2025)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
by: Ma, Chenyang, et al.
Published: (2026)
by: Ma, Chenyang, et al.
Published: (2026)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code
by: Fathallah, Nadeen, et al.
Published: (2025)
by: Fathallah, Nadeen, et al.
Published: (2025)
Fairness Concerns in App Reviews: A Study on AI-based Mobile Apps
by: Nasab, Ali Rezaei, et al.
Published: (2024)
by: Nasab, Ali Rezaei, et al.
Published: (2024)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
by: Wu, Yifan, et al.
Published: (2026)
by: Wu, Yifan, et al.
Published: (2026)
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
by: Gui, Yi, et al.
Published: (2024)
by: Gui, Yi, et al.
Published: (2024)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
by: Zhang, Zejun, et al.
Published: (2025)
by: Zhang, Zejun, et al.
Published: (2025)
Cybernaut: Towards Reliable Web Automation
by: Tomar, Ankur, et al.
Published: (2025)
by: Tomar, Ankur, et al.
Published: (2025)
From PowerPoint UI Sketches to Web-Based Applications: Pattern-Driven Code Generation for GIS Dashboard Development Using Knowledge-Augmented LLMs, Context-Aware Visual Prompting, and the React Framework
by: Xu, Haowen, et al.
Published: (2025)
by: Xu, Haowen, et al.
Published: (2025)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
by: Liang, Qingyuan, et al.
Published: (2021)
by: Liang, Qingyuan, et al.
Published: (2021)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
DetectBERT: Towards Full App-Level Representation Learning to Detect Android Malware
by: Sun, Tiezhu, et al.
Published: (2024)
by: Sun, Tiezhu, et al.
Published: (2024)
Similar Items
-
Insights from Benchmarking Frontier Language Models on Web App Code Generation
by: Cui, Yi
Published: (2024) -
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024) -
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026) -
Rapid Mobile App Development for Generative AI Agents on MIT App Inventor
by: Gao, Jaida, et al.
Published: (2024) -
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)