Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pandit, Shrey, Nguyen, Xuan-Phi, Ming, Yifei, Xu, Austin, Wang, Jiayu, Xiong, Caiming, Joty, Shafiq
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912651174477824
author Pandit, Shrey
Nguyen, Xuan-Phi
Ming, Yifei
Xu, Austin
Wang, Jiayu
Xiong, Caiming
Joty, Shafiq
author_facet Pandit, Shrey
Nguyen, Xuan-Phi
Ming, Yifei
Xu, Austin
Wang, Jiayu
Xiong, Caiming
Joty, Shafiq
contents Web-based 'deep research' agents aim to solve complex question - answering tasks through long-horizon interactions with online tools. These tasks remain challenging, as the underlying language models are often not optimized for long-horizon reasoning and exploration. Prior work has proposed workflows for constructing instruction-tuning datasets, often leveraging knowledge graphs. However, such methods typically lack fine-grained control over difficulty and quality, yielding synthetic data that falls short of capturing the complexity required for long-horizon reasoning. Furthermore, many studies conflate data and training effects by comparing models trained under different optimization recipes, making it difficult to isolate and evaluate the effectiveness of the data itself. We introduce a two-pronged data synthesis pipeline that generates question - answer pairs by progressively increasing task complexity until a frontier baseline web agent fails. The baseline agent plays multiple roles in this process: attempting the questions, validating factuality, checking for alternative answers, and enforcing filtering. To evaluate the effectiveness of our synthesis methods, we adopt a controlled training setup based on distillation from strong web agents. Experiments across multiple web-based benchmarks show that our dataset - despite being smaller - enables the training of more effective web agents than existing datasets. In particular, our data exhibits twice the diversity in tool-use actions, allowing models trained on it to achieve stronger performance while avoiding repetitive tool-calling behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
Pandit, Shrey
Nguyen, Xuan-Phi
Ming, Yifei
Xu, Austin
Wang, Jiayu
Xiong, Caiming
Joty, Shafiq
Computation and Language
Artificial Intelligence
Web-based 'deep research' agents aim to solve complex question - answering tasks through long-horizon interactions with online tools. These tasks remain challenging, as the underlying language models are often not optimized for long-horizon reasoning and exploration. Prior work has proposed workflows for constructing instruction-tuning datasets, often leveraging knowledge graphs. However, such methods typically lack fine-grained control over difficulty and quality, yielding synthetic data that falls short of capturing the complexity required for long-horizon reasoning. Furthermore, many studies conflate data and training effects by comparing models trained under different optimization recipes, making it difficult to isolate and evaluate the effectiveness of the data itself. We introduce a two-pronged data synthesis pipeline that generates question - answer pairs by progressively increasing task complexity until a frontier baseline web agent fails. The baseline agent plays multiple roles in this process: attempting the questions, validating factuality, checking for alternative answers, and enforcing filtering. To evaluate the effectiveness of our synthesis methods, we adopt a controlled training setup based on distillation from strong web agents. Experiments across multiple web-based benchmarks show that our dataset - despite being smaller - enables the training of more effective web agents than existing datasets. In particular, our data exhibits twice the diversity in tool-use actions, allowing models trained on it to achieve stronger performance while avoiding repetitive tool-calling behaviors.
title Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.13913