InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qiyao, Hu, Haoran, Chen, Longze, Wang, Hongbo, Alinejad-Rokny, Hamid, Lin, Yuan, Yang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
by: Wang, Qiyao, et al.
Published: (2026)
by: Wang, Qiyao, et al.
Published: (2026)
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration
by: Wang, Qiyao, et al.
Published: (2026)
by: Wang, Qiyao, et al.
Published: (2026)
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
by: Xu, Ancheng, et al.
Published: (2025)
by: Xu, Ancheng, et al.
Published: (2025)
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
by: Wang, Qiyao, et al.
Published: (2024)
by: Wang, Qiyao, et al.
Published: (2024)
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
by: Chen, Longze, et al.
Published: (2025)
by: Chen, Longze, et al.
Published: (2025)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents
by: Bohra, Arth, et al.
Published: (2025)
by: Bohra, Arth, et al.
Published: (2025)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
by: Zhu, Jingwei, et al.
Published: (2024)
by: Zhu, Jingwei, et al.
Published: (2024)
RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning
by: Chen, Yukun, et al.
Published: (2026)
by: Chen, Yukun, et al.
Published: (2026)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Small Language Model as Data Prospector for Large Language Model
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
by: Li, Shuaimin, et al.
Published: (2025)
by: Li, Shuaimin, et al.
Published: (2025)
Training Superior Sparse Autoencoders for Instruct Models
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property
by: Wang, Qiyao, et al.
Published: (2025)
by: Wang, Qiyao, et al.
Published: (2025)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
by: Chen, Guhong, et al.
Published: (2024)
by: Chen, Guhong, et al.
Published: (2024)
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
by: Chen, Dingwei, et al.
Published: (2024)
by: Chen, Dingwei, et al.
Published: (2024)
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
by: Dai, Dasen, et al.
Published: (2026)
by: Dai, Dasen, et al.
Published: (2026)
PLOT: Enhancing Preference Learning via Optimal Transport
by: Zhu, Liang, et al.
Published: (2026)
by: Zhu, Liang, et al.
Published: (2026)
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025)
by: Li, Renhao, et al.
Published: (2025)
PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation
by: Luo, Jing, et al.
Published: (2024)
by: Luo, Jing, et al.
Published: (2024)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
by: Chen, Dingwei, et al.
Published: (2025)
by: Chen, Dingwei, et al.
Published: (2025)
Safe and Scalable Web Agent Learning via Recreated Websites
by: Chae, Hyungjoo, et al.
Published: (2026)
by: Chae, Hyungjoo, et al.
Published: (2026)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
ExecRepoBench: Multi-level Executable Code Completion Evaluation
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
by: Yoran, Ori, et al.
Published: (2024)
by: Yoran, Ori, et al.
Published: (2024)
Modeling Distinct Human Interaction in Web Agents
by: Huq, Faria, et al.
Published: (2026)
by: Huq, Faria, et al.
Published: (2026)
Learning Ordinal Probabilistic Reward from Preferences
by: Chen, Longze, et al.
Published: (2026)
by: Chen, Longze, et al.
Published: (2026)
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
GeoBuildBench: A Benchmark for Interactive and Executable Geometry Construction from Natural Language
by: Kim, Jinwoong, et al.
Published: (2026)
by: Kim, Jinwoong, et al.
Published: (2026)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution
by: Lin, Leon, et al.
Published: (2025)
by: Lin, Leon, et al.
Published: (2025)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
by: Li, Siyi, et al.
Published: (2026)
by: Li, Siyi, et al.
Published: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions
by: Liang, Yuheng, et al.
Published: (2024)
by: Liang, Yuheng, et al.
Published: (2024)
Similar Items
-
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
by: Wang, Qiyao, et al.
Published: (2026) -
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration
by: Wang, Qiyao, et al.
Published: (2026) -
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
by: Xu, Ancheng, et al.
Published: (2025) -
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
by: Wang, Qiyao, et al.
Published: (2024) -
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
by: Chen, Longze, et al.
Published: (2025)