WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Zimu, Ren, Houxing, Yang, Yunqiao, Wang, Ke, Zong, Zhuofan, Pan, Junting, Zhan, Mingjie, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026)
by: Lu, Zimu, et al.
Published: (2026)
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
by: Yang, Yunqiao, et al.
Published: (2026)
by: Yang, Yunqiao, et al.
Published: (2026)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
Edit-Based Refinement for Parallel Masked Diffusion Language Models
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
by: Yang, Yunqiao, et al.
Published: (2025)
by: Yang, Yunqiao, et al.
Published: (2025)
Alignment with Fill-In-the-Middle for Enhancing Code Generation
by: Ren, Houxing, et al.
Published: (2025)
by: Ren, Houxing, et al.
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
by: Shi, Weikang, et al.
Published: (2026)
by: Shi, Weikang, et al.
Published: (2026)
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning
by: Jiang, Juyong, et al.
Published: (2026)
by: Jiang, Juyong, et al.
Published: (2026)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation
by: Wang, Kuang-Da, et al.
Published: (2025)
by: Wang, Kuang-Da, et al.
Published: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
by: Huang, Zhiyuan, et al.
Published: (2025)
by: Huang, Zhiyuan, et al.
Published: (2025)
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
Full‐Range Adaptive Charging Control for PT‐Symmetric Wireless Power Transfer Systems Under Variable Load and Mutual Inductance
by: Zekai Wang, et al.
Published: (2025)
by: Zekai Wang, et al.
Published: (2025)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
by: Li, Junsong, et al.
Published: (2025)
by: Li, Junsong, et al.
Published: (2025)
Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
by: Ma, Bingqi, et al.
Published: (2024)
by: Ma, Bingqi, et al.
Published: (2024)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
by: Wang, Qiyao, et al.
Published: (2026)
by: Wang, Qiyao, et al.
Published: (2026)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
by: Xiong, Weimin, et al.
Published: (2024)
by: Xiong, Weimin, et al.
Published: (2024)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning
by: Liu, Xirui, et al.
Published: (2026)
by: Liu, Xirui, et al.
Published: (2026)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
by: He, Zehai, et al.
Published: (2026)
by: He, Zehai, et al.
Published: (2026)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
by: Shi, Weikang, et al.
Published: (2025)
by: Shi, Weikang, et al.
Published: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Safe and Scalable Web Agent Learning via Recreated Websites
by: Chae, Hyungjoo, et al.
Published: (2026)
by: Chae, Hyungjoo, et al.
Published: (2026)
WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents
by: Bohra, Arth, et al.
Published: (2025)
by: Bohra, Arth, et al.
Published: (2025)
Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
RepoGenReflex: Enhancing Repository-Level Code Completion with Verbal Reinforcement and Retrieval-Augmented Generation
by: Wang, Jicheng, et al.
Published: (2024)
by: Wang, Jicheng, et al.
Published: (2024)
LLM Agents can Autonomously Hack Websites
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
by: Shao, Hao, et al.
Published: (2026)
by: Shao, Hao, et al.
Published: (2026)
Similar Items
-
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025) -
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026) -
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
by: Yang, Yunqiao, et al.
Published: (2026) -
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
by: Ren, Houxing, et al.
Published: (2026) -
Edit-Based Refinement for Parallel Masked Diffusion Language Models
by: Ren, Houxing, et al.
Published: (2026)