WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Awal, Rabiul, Massoud, Mahsa, Feizi, Aarash, Li, Zichao, Wang, Suyuchen, Pal, Christopher, Agrawal, Aishwarya, Vazquez, David, Reddy, Siva, Rodriguez, Juan A., Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908499990020096
author Awal, Rabiul
Massoud, Mahsa
Feizi, Aarash
Li, Zichao
Wang, Suyuchen
Pal, Christopher
Agrawal, Aishwarya
Vazquez, David
Reddy, Siva
Rodriguez, Juan A.
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
author_facet Awal, Rabiul
Massoud, Mahsa
Feizi, Aarash
Li, Zichao
Wang, Suyuchen
Pal, Christopher
Agrawal, Aishwarya
Vazquez, David
Reddy, Siva
Rodriguez, Juan A.
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
contents We present WebMMU, a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. Unlike prior benchmarks that treat these tasks separately, WebMMU unifies them using expert-annotated, real-world web data to assess models' abilities in complex multi-step reasoning, precise element grounding, and functional UI comprehension and coding. Our evaluation shows that while multimodal large language models (MLLMs) perform well on basic information extraction, they struggle with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. These findings reveal key limitations in current MLLMs and underscore the need for improved multimodal and cross-lingual reasoning to build future web agents capable of automating diverse web development tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
Awal, Rabiul
Massoud, Mahsa
Feizi, Aarash
Li, Zichao
Wang, Suyuchen
Pal, Christopher
Agrawal, Aishwarya
Vazquez, David
Reddy, Siva
Rodriguez, Juan A.
Taslakian, Perouz
Gella, Spandana
Rajeswar, Sai
Computer Vision and Pattern Recognition
We present WebMMU, a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. Unlike prior benchmarks that treat these tasks separately, WebMMU unifies them using expert-annotated, real-world web data to assess models' abilities in complex multi-step reasoning, precise element grounding, and functional UI comprehension and coding. Our evaluation shows that while multimodal large language models (MLLMs) perform well on basic information extraction, they struggle with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. These findings reveal key limitations in current MLLMs and underscore the need for improved multimodal and cross-lingual reasoning to build future web agents capable of automating diverse web development tasks.
title WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.16763