Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Aiden Yiliu, Hao, Xinyue, Liu, Shilong, Wang, Mengdi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911416709021696
author Li, Aiden Yiliu
Hao, Xinyue
Liu, Shilong
Wang, Mengdi
author_facet Li, Aiden Yiliu
Hao, Xinyue
Liu, Shilong
Wang, Mengdi
contents Despite advances in multimodal large language models, autonomous web agents still struggle to reliably execute long-horizon tasks on complex and dynamic web interfaces. Existing agents often suffer from inaccurate element grounding, the absence of site-specific procedural knowledge, and unstable long-term task tracking and memory, particularly when operating over complex Document Object Model structures. To address these limitations, we introduce Avenir-Web, a web agent that achieves a new open-source state of the art on the Online-Mind2Web benchmark in real-world deployment. Avenir-Web leverages a Mixture of Grounding Experts, Experience-Imitation Planning for incorporating procedural priors, and a task-tracking checklist combined with adaptive memory to enable robust and seamless interaction across diverse user interface paradigms. We evaluate Avenir-Web on Online-Mind2Web, a rigorous benchmark of live and user-centered web tasks. Our results demonstrate that Avenir-Web significantly surpasses prior open-source agents and attains performance parity with top-tier proprietary models, thereby establishing a new open-source state of the art for reliable web agents on live websites.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02468
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
Li, Aiden Yiliu
Hao, Xinyue
Liu, Shilong
Wang, Mengdi
Artificial Intelligence
Computation and Language
Despite advances in multimodal large language models, autonomous web agents still struggle to reliably execute long-horizon tasks on complex and dynamic web interfaces. Existing agents often suffer from inaccurate element grounding, the absence of site-specific procedural knowledge, and unstable long-term task tracking and memory, particularly when operating over complex Document Object Model structures. To address these limitations, we introduce Avenir-Web, a web agent that achieves a new open-source state of the art on the Online-Mind2Web benchmark in real-world deployment. Avenir-Web leverages a Mixture of Grounding Experts, Experience-Imitation Planning for incorporating procedural priors, and a task-tracking checklist combined with adaptive memory to enable robust and seamless interaction across diverse user interface paradigms. We evaluate Avenir-Web on Online-Mind2Web, a rigorous benchmark of live and user-centered web tasks. Our results demonstrate that Avenir-Web significantly surpasses prior open-source agents and attains performance parity with top-tier proprietary models, thereby establishing a new open-source state of the art for reliable web agents on live websites.
title Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.02468