Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Junliang, Xiao, Jingyu, Tang, Wenxin, Wang, Zhixian, Xie, Zipeng, Wang, Wenxuan, Zhang, Minrui, Yu, Shuanghe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
by: Xiao, Jingyu, et al.
Published: (2024)
by: Xiao, Jingyu, et al.
Published: (2024)
SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design
by: Tang, Wenxin, et al.
Published: (2025)
by: Tang, Wenxin, et al.
Published: (2025)
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
by: Wen, Ming, et al.
Published: (2026)
by: Wen, Ming, et al.
Published: (2026)
EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection
by: Tang, Wenxin, et al.
Published: (2026)
by: Tang, Wenxin, et al.
Published: (2026)
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
by: Xue, Qi, et al.
Published: (2025)
by: Xue, Qi, et al.
Published: (2025)
VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection
by: Tang, Wenxin, et al.
Published: (2026)
by: Tang, Wenxin, et al.
Published: (2026)
EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
by: Guo, Haiyun, et al.
Published: (2025)
by: Guo, Haiyun, et al.
Published: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization
by: Tang, Wenxin, et al.
Published: (2026)
by: Tang, Wenxin, et al.
Published: (2026)
Internalizing Safety Understanding in Large Reasoning Models via Verification
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Benchmarking Reasoning Robustness in Large Language Models
by: Yu, Tong, et al.
Published: (2025)
by: Yu, Tong, et al.
Published: (2025)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
by: Lu, Jinghui, et al.
Published: (2025)
by: Lu, Jinghui, et al.
Published: (2025)
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
by: Shi, Yichen, et al.
Published: (2025)
by: Shi, Yichen, et al.
Published: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
by: Xiong, Chuyan, et al.
Published: (2024)
by: Xiong, Chuyan, et al.
Published: (2024)
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities
by: Zeng, Changyu, et al.
Published: (2025)
by: Zeng, Changyu, et al.
Published: (2025)
Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
by: Tang, Jiaqi, et al.
Published: (2025)
by: Tang, Jiaqi, et al.
Published: (2025)
Towards Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-critical Scenarios
by: Li, Jingzheng, et al.
Published: (2025)
by: Li, Jingzheng, et al.
Published: (2025)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
by: Levy, Ido, et al.
Published: (2024)
by: Levy, Ido, et al.
Published: (2024)
Reinforced Latent Reasoning for LLM-based Recommendation
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
by: Mahesheka, Harsh, et al.
Published: (2024)
by: Mahesheka, Harsh, et al.
Published: (2024)
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs
by: Ning, Yansong, et al.
Published: (2026)
by: Ning, Yansong, et al.
Published: (2026)
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
by: Song, Zhixue, et al.
Published: (2026)
by: Song, Zhixue, et al.
Published: (2026)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
by: Qi, Haozhe, et al.
Published: (2026)
by: Qi, Haozhe, et al.
Published: (2026)
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
by: Zhang, Haoyuan, et al.
Published: (2026)
by: Zhang, Haoyuan, et al.
Published: (2026)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026)
by: Bei, Yuanchen, et al.
Published: (2026)
WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
by: Zhou, Guanghao, et al.
Published: (2025)
by: Zhou, Guanghao, et al.
Published: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026)
by: Bai, Haoyue, et al.
Published: (2026)
WebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web Agent
by: Zhang, Lingfeng, et al.
Published: (2026)
by: Zhang, Lingfeng, et al.
Published: (2026)
Similar Items
-
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025) -
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
by: Xiao, Jingyu, et al.
Published: (2024) -
SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design
by: Tang, Wenxin, et al.
Published: (2025) -
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
by: Wen, Ming, et al.
Published: (2026) -
EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection
by: Tang, Wenxin, et al.
Published: (2026)