Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
Fuente:
arXiv
Salvato in:
| Autori principali: | Tran, Hung, Nashold, Langston, Krishnan, Rayan, Bigeard, Antoine, Gu, Alex |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
di: Young, Richard J.
Pubblicazione: (2025)
di: Young, Richard J.
Pubblicazione: (2025)
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025)
di: Miller, Samuel, et al.
Pubblicazione: (2025)
CIFE: Code Instruction-Following Evaluation
di: Gunnu, Sravani, et al.
Pubblicazione: (2025)
di: Gunnu, Sravani, et al.
Pubblicazione: (2025)
Simple and Effective Baselines for Code Summarisation Evaluation
di: Robinson, Jade, et al.
Pubblicazione: (2025)
di: Robinson, Jade, et al.
Pubblicazione: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
Engineering A Large Language Model From Scratch
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2024)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2024)
Can AI Assist in Olympiad Coding
di: Ren, Samuel
Pubblicazione: (2025)
di: Ren, Samuel
Pubblicazione: (2025)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
di: Thorne, Simon, et al.
Pubblicazione: (2025)
di: Thorne, Simon, et al.
Pubblicazione: (2025)
Plan with Code: Comparing approaches for robust NL to DSL generation
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
Enhancing LLM Code Generation Capabilities through Test-Driven Development and Code Interpreter
di: Jalil, Sajed, et al.
Pubblicazione: (2025)
di: Jalil, Sajed, et al.
Pubblicazione: (2025)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
di: He, Kaifeng, et al.
Pubblicazione: (2025)
di: He, Kaifeng, et al.
Pubblicazione: (2025)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
di: Bassamzadeh, Nastaran, et al.
Pubblicazione: (2024)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
di: Bigeard, Antoine, et al.
Pubblicazione: (2025)
di: Bigeard, Antoine, et al.
Pubblicazione: (2025)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
di: Weng, Haojun, et al.
Pubblicazione: (2026)
di: Weng, Haojun, et al.
Pubblicazione: (2026)
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
di: Alam, Khairul, et al.
Pubblicazione: (2024)
di: Alam, Khairul, et al.
Pubblicazione: (2024)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
di: Rai, Daking, et al.
Pubblicazione: (2025)
di: Rai, Daking, et al.
Pubblicazione: (2025)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
di: Zhu, Andrew, et al.
Pubblicazione: (2024)
di: Zhu, Andrew, et al.
Pubblicazione: (2024)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
di: Sakizli, Furkan
Pubblicazione: (2026)
di: Sakizli, Furkan
Pubblicazione: (2026)
MicroRemed: Benchmarking LLMs in Microservices Remediation
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
Narrow Transformer: StarCoder-Based Java-LM For Desktop
di: Rathinasamy, Kamalkumar, et al.
Pubblicazione: (2024)
di: Rathinasamy, Kamalkumar, et al.
Pubblicazione: (2024)
ContractBench: Can LLM Agents Preserve Observation Contracts?
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
di: Wang, Jicheng, et al.
Pubblicazione: (2026)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
di: Abualazm, Raafat, et al.
Pubblicazione: (2026)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
di: Mughal, Ali Hassaan, et al.
Pubblicazione: (2026)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
di: Karpurapu, Shanthi, et al.
Pubblicazione: (2024)
di: Karpurapu, Shanthi, et al.
Pubblicazione: (2024)
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
di: Gringras, David
Pubblicazione: (2026)
di: Gringras, David
Pubblicazione: (2026)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
di: Lim, Soohan, et al.
Pubblicazione: (2025)
di: Lim, Soohan, et al.
Pubblicazione: (2025)
Learning Software Bug Reports: A Systematic Literature Review
di: Long, Guoming, et al.
Pubblicazione: (2025)
di: Long, Guoming, et al.
Pubblicazione: (2025)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
di: Li, Caihua, et al.
Pubblicazione: (2026)
di: Li, Caihua, et al.
Pubblicazione: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
di: Li, Kenan, et al.
Pubblicazione: (2026)
di: Li, Kenan, et al.
Pubblicazione: (2026)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
di: Mazaheri, Parsa
Pubblicazione: (2026)
di: Mazaheri, Parsa
Pubblicazione: (2026)
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
di: Savenkov, Vladislav
Pubblicazione: (2026)
di: Savenkov, Vladislav
Pubblicazione: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
di: Nguyen, Quang-Dung, et al.
Pubblicazione: (2025)
di: Nguyen, Quang-Dung, et al.
Pubblicazione: (2025)
Anka: A Domain-Specific Language for Reliable LLM Code Generation
di: Mazrouei, Saif Khalfan Saif Al
Pubblicazione: (2025)
di: Mazrouei, Saif Khalfan Saif Al
Pubblicazione: (2025)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
di: Kohl, Jens, et al.
Pubblicazione: (2024)
di: Kohl, Jens, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025) -
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
di: Young, Richard J.
Pubblicazione: (2025) -
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025) -
CIFE: Code Instruction-Following Evaluation
di: Gunnu, Sravani, et al.
Pubblicazione: (2025) -
Simple and Effective Baselines for Code Summarisation Evaluation
di: Robinson, Jade, et al.
Pubblicazione: (2025)