LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoang, Thai, Huang, Kung-Hsiang, Kokane, Shirley, Zhang, Jianguo, Liu, Zuxin, Zhu, Ming, Grigsby, Jake, Lan, Tian, Ryoo, Michael S, Wu, Chien-Sheng, Heinecke, Shelby, Wang, Huan, Savarese, Silvio, Xiong, Caiming, Niebles, Juan Carlos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913872072409088
author Hoang, Thai
Huang, Kung-Hsiang
Kokane, Shirley
Zhang, Jianguo
Liu, Zuxin
Zhu, Ming
Grigsby, Jake
Lan, Tian
Ryoo, Michael S
Wu, Chien-Sheng
Heinecke, Shelby
Wang, Huan
Savarese, Silvio
Xiong, Caiming
Niebles, Juan Carlos
author_facet Hoang, Thai
Huang, Kung-Hsiang
Kokane, Shirley
Zhang, Jianguo
Liu, Zuxin
Zhu, Ming
Grigsby, Jake
Lan, Tian
Ryoo, Michael S
Wu, Chien-Sheng
Heinecke, Shelby
Wang, Huan
Savarese, Silvio
Xiong, Caiming
Niebles, Juan Carlos
contents Large Action Models (LAMs) for AI Agents offer incredible potential but face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. To address these issues, we present LAM SIMULATOR, a comprehensive framework designed for online exploration of agentic tasks with high-quality feedback. Our framework features a dynamic task query generator, an extensive collection of tools, and an interactive environment where Large Language Model (LLM) Agents can call tools and receive real-time feedback. This setup enables LLM Agents to explore and solve tasks autonomously, facilitating the discovery of multiple approaches to tackle any given task. The resulting action trajectory data are then used to create high-quality training datasets for LAMs. Our experiments on popular agentic benchmarks, ToolBench and CRMArena, highlight the effectiveness of LAM SIMULATOR: models trained with self-generated datasets using our framework achieve significant performance gains, up to a 49.3\% improvement over their original baselines. LAM SIMULATOR requires minimal human input during dataset creation, highlighting LAM SIMULATOR's efficiency and effectiveness in speeding up development of AI agents.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
Hoang, Thai
Huang, Kung-Hsiang
Kokane, Shirley
Zhang, Jianguo
Liu, Zuxin
Zhu, Ming
Grigsby, Jake
Lan, Tian
Ryoo, Michael S
Wu, Chien-Sheng
Heinecke, Shelby
Wang, Huan
Savarese, Silvio
Xiong, Caiming
Niebles, Juan Carlos
Computation and Language
Artificial Intelligence
Machine Learning
Large Action Models (LAMs) for AI Agents offer incredible potential but face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. To address these issues, we present LAM SIMULATOR, a comprehensive framework designed for online exploration of agentic tasks with high-quality feedback. Our framework features a dynamic task query generator, an extensive collection of tools, and an interactive environment where Large Language Model (LLM) Agents can call tools and receive real-time feedback. This setup enables LLM Agents to explore and solve tasks autonomously, facilitating the discovery of multiple approaches to tackle any given task. The resulting action trajectory data are then used to create high-quality training datasets for LAMs. Our experiments on popular agentic benchmarks, ToolBench and CRMArena, highlight the effectiveness of LAM SIMULATOR: models trained with self-generated datasets using our framework achieve significant performance gains, up to a 49.3\% improvement over their original baselines. LAM SIMULATOR requires minimal human input during dataset creation, highlighting LAM SIMULATOR's efficiency and effectiveness in speeding up development of AI agents.
title LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.02298