WebApp1K: A Practical Code-Generation Benchmark for Web App Development

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Cui, Yi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911974117343232
author Cui, Yi
author_facet Cui, Yi
contents We introduce WebApp1K, a practical code-generation benchmark to measure LLM ability to develop web apps. This benchmark aims to calibrate LLM output and aid the models to progressively improve code correctness and functionality. The benchmark is lightweight and easy to run. We present the initial version of WebApp1K, and share our findings of running the benchmark against the latest frontier LLMs. First, open source LLMs deliver impressive performance, closely trailing behind GPT-4o and Claude 3.5. Second, model size has strong correlation with code correctness. Third, no prompting techniques have been found to lift performance either universally to all models, or significantly to a single model.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00019
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WebApp1K: A Practical Code-Generation Benchmark for Web App Development
Cui, Yi
Software Engineering
Artificial Intelligence
We introduce WebApp1K, a practical code-generation benchmark to measure LLM ability to develop web apps. This benchmark aims to calibrate LLM output and aid the models to progressively improve code correctness and functionality. The benchmark is lightweight and easy to run. We present the initial version of WebApp1K, and share our findings of running the benchmark against the latest frontier LLMs. First, open source LLMs deliver impressive performance, closely trailing behind GPT-4o and Claude 3.5. Second, model size has strong correlation with code correctness. Third, no prompting techniques have been found to lift performance either universally to all models, or significantly to a single model.
title WebApp1K: A Practical Code-Generation Benchmark for Web App Development
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2408.00019