Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Yun, Sukmin, Lin, Haokun, Thushara, Rusiru, Bhat, Mohammad Qazim, Wang, Yongxin, Jiang, Zutao, Deng, Mingkai, Wang, Jinhong, Tao, Tianhua, Li, Junbo, Li, Haonan, Nakov, Preslav, Baldwin, Timothy, Liu, Zhengzhong, Xing, Eric P., Liang, Xiaodan, Shen, Zhiqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
di: Orel, Daniil, et al.
Pubblicazione: (2025)
di: Orel, Daniil, et al.
Pubblicazione: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
di: Tao, Tianhua, et al.
Pubblicazione: (2024)
di: Tao, Tianhua, et al.
Pubblicazione: (2024)
Arabic Dataset for LLM Safeguard Evaluation
di: Ashraf, Yasser, et al.
Pubblicazione: (2024)
di: Ashraf, Yasser, et al.
Pubblicazione: (2024)
Detecting Propaganda Techniques in Code-Switched Social Media Text
di: Salman, Muhammad Umar, et al.
Pubblicazione: (2023)
di: Salman, Muhammad Umar, et al.
Pubblicazione: (2023)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
di: Orel, Daniil, et al.
Pubblicazione: (2025)
di: Orel, Daniil, et al.
Pubblicazione: (2025)
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
di: Gui, Yi, et al.
Pubblicazione: (2024)
di: Gui, Yi, et al.
Pubblicazione: (2024)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
di: Thushara, Rusiru, et al.
Pubblicazione: (2026)
di: Thushara, Rusiru, et al.
Pubblicazione: (2026)
Rethinking STS and NLI in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
Loki: An Open-Source Tool for Fact Verification
di: Li, Haonan, et al.
Pubblicazione: (2024)
di: Li, Haonan, et al.
Pubblicazione: (2024)
FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals
di: Yotkova, Elitsa, et al.
Pubblicazione: (2026)
di: Yotkova, Elitsa, et al.
Pubblicazione: (2026)
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
di: Xiao, Jingyu, et al.
Pubblicazione: (2024)
di: Xiao, Jingyu, et al.
Pubblicazione: (2024)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
di: Tomar, Raj Vardhan, et al.
Pubblicazione: (2026)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
di: Alhindi, Tariq, et al.
Pubblicazione: (2023)
di: Alhindi, Tariq, et al.
Pubblicazione: (2023)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
di: Elozeiri, Kareem, et al.
Pubblicazione: (2025)
SimuScene: Training and Benchmarking Code Generation to Simulate Physical Scenarios
di: Wang, Yanan, et al.
Pubblicazione: (2026)
di: Wang, Yanan, et al.
Pubblicazione: (2026)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
di: Das, Rocktim Jyoti, et al.
Pubblicazione: (2024)
di: Das, Rocktim Jyoti, et al.
Pubblicazione: (2024)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
di: Almheiri, Saeed, et al.
Pubblicazione: (2025)
Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformation
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
di: Sundriyal, Megha, et al.
Pubblicazione: (2023)
Adapting Fake News Detection to the Era of Large Language Models
di: Su, Jinyan, et al.
Pubblicazione: (2023)
di: Su, Jinyan, et al.
Pubblicazione: (2023)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
di: Agarwal, Siddhant, et al.
Pubblicazione: (2024)
di: Agarwal, Siddhant, et al.
Pubblicazione: (2024)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
di: Liu, Zhihong, et al.
Pubblicazione: (2026)
di: Liu, Zhihong, et al.
Pubblicazione: (2026)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
di: Mekky, Ali, et al.
Pubblicazione: (2025)
di: Mekky, Ali, et al.
Pubblicazione: (2025)
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks
di: Rugina, Ileana, et al.
Pubblicazione: (2020)
di: Rugina, Ileana, et al.
Pubblicazione: (2020)
Exploring Language Model Generalization in Low-Resource Extractive QA
di: Sengupta, Saptarshi, et al.
Pubblicazione: (2024)
di: Sengupta, Saptarshi, et al.
Pubblicazione: (2024)
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
di: Fu, Zihang, et al.
Pubblicazione: (2026)
di: Fu, Zihang, et al.
Pubblicazione: (2026)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
di: Bates, Luke, et al.
Pubblicazione: (2025)
di: Bates, Luke, et al.
Pubblicazione: (2025)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
di: Vashurin, Roman, et al.
Pubblicazione: (2025)
di: Vashurin, Roman, et al.
Pubblicazione: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
di: Glockner, Max, et al.
Pubblicazione: (2024)
di: Glockner, Max, et al.
Pubblicazione: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
di: Glockner, Max, et al.
Pubblicazione: (2024)
di: Glockner, Max, et al.
Pubblicazione: (2024)
An Analytical Emotion Framework of Rumour Threads on Social Media
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
di: Baliah, Sanoojan, et al.
Pubblicazione: (2026)
di: Baliah, Sanoojan, et al.
Pubblicazione: (2026)
Multimodal Large Language Models to Support Real-World Fact-Checking
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
di: Su, Jinyan, et al.
Pubblicazione: (2025)
di: Su, Jinyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
di: Orel, Daniil, et al.
Pubblicazione: (2025) -
Crystal: Illuminating LLM Abilities on Language and Code
di: Tao, Tianhua, et al.
Pubblicazione: (2024) -
Arabic Dataset for LLM Safeguard Evaluation
di: Ashraf, Yasser, et al.
Pubblicazione: (2024) -
Detecting Propaganda Techniques in Code-Switched Social Media Text
di: Salman, Muhammad Umar, et al.
Pubblicazione: (2023) -
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
di: Orel, Daniil, et al.
Pubblicazione: (2025)