ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coca, Alexandru, Gaynor, Mark, Zhang, Zhenxing, Cheng, Jianpeng, Tseng, Bo-Hsiang, Boothroyd, Pete, Alonso, Héctor Martinez, Séaghdha, Diarmuid Ó, Johannsen, Anders
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913950965170176
author Coca, Alexandru
Gaynor, Mark
Zhang, Zhenxing
Cheng, Jianpeng
Tseng, Bo-Hsiang
Boothroyd, Pete
Alonso, Héctor Martinez
Séaghdha, Diarmuid Ó
Johannsen, Anders
author_facet Coca, Alexandru
Gaynor, Mark
Zhang, Zhenxing
Cheng, Jianpeng
Tseng, Bo-Hsiang
Boothroyd, Pete
Alonso, Héctor Martinez
Séaghdha, Diarmuid Ó
Johannsen, Anders
contents This work evaluates the potential of large language models (LLMs) to power digital assistants capable of complex action execution. These assistants rely on pre-trained programming knowledge to execute multi-step goals by composing objects and functions defined in assistant libraries into action execution programs. To achieve this, we develop ASPERA, a framework comprising an assistant library simulation and a human-assisted LLM data generation engine. Our engine allows developers to guide LLM generation of high-quality tasks consisting of complex user queries, simulation state and corresponding validation programs, tackling data availability and evaluation robustness challenges. Alongside the framework we release Asper-Bench, an evaluation dataset of 250 challenging tasks generated using ASPERA, which we use to show that program generation grounded in custom assistant libraries is a significant challenge to LLMs compared to dependency-free code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
Coca, Alexandru
Gaynor, Mark
Zhang, Zhenxing
Cheng, Jianpeng
Tseng, Bo-Hsiang
Boothroyd, Pete
Alonso, Héctor Martinez
Séaghdha, Diarmuid Ó
Johannsen, Anders
Computation and Language
Artificial Intelligence
Machine Learning
This work evaluates the potential of large language models (LLMs) to power digital assistants capable of complex action execution. These assistants rely on pre-trained programming knowledge to execute multi-step goals by composing objects and functions defined in assistant libraries into action execution programs. To achieve this, we develop ASPERA, a framework comprising an assistant library simulation and a human-assisted LLM data generation engine. Our engine allows developers to guide LLM generation of high-quality tasks consisting of complex user queries, simulation state and corresponding validation programs, tackling data availability and evaluation robustness challenges. Alongside the framework we release Asper-Bench, an evaluation dataset of 250 challenging tasks generated using ASPERA, which we use to show that program generation grounded in custom assistant libraries is a significant challenge to LLMs compared to dependency-free code generation.
title ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.15501