WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chunyang, Zheng, Yilun, Huang, Xinting, Fang, Tianqing, Xu, Jiahao, Chen, Lihui, Song, Yangqiu, Hu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
A Web-Based IDE for DevOps Learning in Software Engineering Higher Education
by: Iyer, Ganesh Neelakanta, et al.
Published: (2024)
by: Iyer, Ganesh Neelakanta, et al.
Published: (2024)
DevMuT: Testing Deep Learning Framework via Developer Expertise-Based Mutation
by: Mu, Yanzhou, et al.
Published: (2025)
by: Mu, Yanzhou, et al.
Published: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
by: Du, Yaxin, et al.
Published: (2025)
by: Du, Yaxin, et al.
Published: (2025)
Embedded DevOps: A Survey on the Application of DevOps Practices in Embedded Software and Firmware Development
by: Katapara, Parthiv, et al.
Published: (2025)
by: Katapara, Parthiv, et al.
Published: (2025)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
by: Lei, Xinping, et al.
Published: (2026)
by: Lei, Xinping, et al.
Published: (2026)
DevGPT: Studying Developer-ChatGPT Conversations
by: Xiao, Tao, et al.
Published: (2023)
by: Xiao, Tao, et al.
Published: (2023)
SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
WebMAC: A Multi-Agent Collaborative Framework for Scenario Testing of Web Systems
by: Wan, Zhenyu, et al.
Published: (2026)
by: Wan, Zhenyu, et al.
Published: (2026)
Towards a Science of Developer eXperience (DevX)
by: Combemale, Benoit
Published: (2025)
by: Combemale, Benoit
Published: (2025)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
Development of an Automated Web Application for Efficient Web Scraping: Design and Implementation
by: Dutta, Alok, et al.
Published: (2025)
by: Dutta, Alok, et al.
Published: (2025)
A Survey on Web Application Testing: A Decade of Evolution
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
ChatDev: Communicative Agents for Software Development
by: Qian, Chen, et al.
Published: (2023)
by: Qian, Chen, et al.
Published: (2023)
WEFix: Intelligent Automatic Generation of Explicit Waits for Efficient Web End-to-End Flaky Tests
by: Liu, Xinyue, et al.
Published: (2024)
by: Liu, Xinyue, et al.
Published: (2024)
iReDev: A Knowledge-Driven Multi-Agent Framework for Intelligent Requirements Development
by: Jin, Dongming, et al.
Published: (2025)
by: Jin, Dongming, et al.
Published: (2025)
WebSPL: A Software Product Line for Web Applications
by: da Luz, Maicon Azevedo, et al.
Published: (2024)
by: da Luz, Maicon Azevedo, et al.
Published: (2024)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
by: Peters, Gideon, et al.
Published: (2026)
by: Peters, Gideon, et al.
Published: (2026)
TESTQUEST: A Web Gamification Tool to Improve Locators and Page Objects Quality
by: Olianas, Dario, et al.
Published: (2025)
by: Olianas, Dario, et al.
Published: (2025)
AutoDev: Automated AI-Driven Development
by: Tufano, Michele, et al.
Published: (2024)
by: Tufano, Michele, et al.
Published: (2024)
The road to Sustainable DevOps
by: Herati, Darwish Ahmad, et al.
Published: (2025)
by: Herati, Darwish Ahmad, et al.
Published: (2025)
Chat-like Asserts Prediction with the Support of Large Language Model
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Envisioning Future Interactive Web Development: Editing Webpage with Natural Language
by: Dang, Truong Hai, et al.
Published: (2025)
by: Dang, Truong Hai, et al.
Published: (2025)
A Prototype VS Code Extension to Improve Web Accessible Development
by: Calì, Elisa, et al.
Published: (2025)
by: Calì, Elisa, et al.
Published: (2025)
Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
by: Huang, Kai, et al.
Published: (2025)
by: Huang, Kai, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
by: Teoh, Xiwen, et al.
Published: (2026)
by: Teoh, Xiwen, et al.
Published: (2026)
Research on WebAssembly Runtimes: A Survey
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
APITestGenie: Generating Web API Tests from Requirements and API Specifications with LLMs
by: Pereira, André, et al.
Published: (2026)
by: Pereira, André, et al.
Published: (2026)
Web Element Relocalization in Evolving Web Applications: A Comparative Analysis and Extension Study
by: Kluge, Anton, et al.
Published: (2025)
by: Kluge, Anton, et al.
Published: (2025)
DevServOps: DevOps For Product-Oriented Product-Service Systems
by: Dakkak, Anas, et al.
Published: (2023)
by: Dakkak, Anas, et al.
Published: (2023)
USeR: A Web-based User Story eReviewer for Assisted Quality Optimizations
by: Hallmann, Daniel, et al.
Published: (2025)
by: Hallmann, Daniel, et al.
Published: (2025)
Design and Development of a Web Platform for Blood Donation Management
by: Ali, Fatima Zulfiqar, et al.
Published: (2025)
by: Ali, Fatima Zulfiqar, et al.
Published: (2025)
Neural Embeddings for Web Testing
by: Kanaththage, Kasun, et al.
Published: (2023)
by: Kanaththage, Kasun, et al.
Published: (2023)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases
by: Yu, Gerhard, et al.
Published: (2025)
by: Yu, Gerhard, et al.
Published: (2025)
Similar Items
-
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026) -
A Web-Based IDE for DevOps Learning in Software Engineering Higher Education
by: Iyer, Ganesh Neelakanta, et al.
Published: (2024) -
DevMuT: Testing Deep Learning Framework via Developer Expertise-Based Mutation
by: Mu, Yanzhou, et al.
Published: (2025) -
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024) -
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)