Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Jingxing, Zhou, Chenyu, Fu, Zhihui, Wang, Jun, Liu, Weiwen, Zhang, Weinan, Lin, Jianghao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910226382323712
author Wang, Jingxing
Zhou, Chenyu
Fu, Zhihui
Wang, Jun
Liu, Weiwen
Zhang, Weinan
Lin, Jianghao
author_facet Wang, Jingxing
Zhou, Chenyu
Fu, Zhihui
Wang, Jun
Liu, Weiwen
Zhang, Weinan
Lin, Jianghao
contents LLM agents benefit from reusable skills, yet test-time tasks often require guidance more specific than a static skill library can provide. We propose \emph{SkillTTA}, a Test-Time Adaptive Skill Synthesis method that retrieves a small set of training trajectories relevant to the current task and synthesizes them into a temporary, task-specific textual skill. The solver model is kept fixed, so adaptation happens entirely through generated context rather than parameter updates. We evaluate the method on SpreadsheetBench, ALFWorld, and BigCodeBench. Compared with static trajectory-to-skill synthesis using GPT-5.5, task-specific skills improve SpreadsheetBench Pass@1 from 0.397 to 0.505 and BigCodeBench Pass@1 from 0.517 to 0.651. On ALFWorld, the method matches a heavier memory-learning baseline within four points of success rate while producing the shortest successful trajectories among reported methods. Ablations on SpreadsheetBench further show that synthesized skills outperform raw trajectory prompting, that top-$k$ retrieval should stay small, and that failed trajectories are especially useful because they expose recurring evaluator-facing mistakes.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16986
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
Wang, Jingxing
Zhou, Chenyu
Fu, Zhihui
Wang, Jun
Liu, Weiwen
Zhang, Weinan
Lin, Jianghao
Computation and Language
Artificial Intelligence
LLM agents benefit from reusable skills, yet test-time tasks often require guidance more specific than a static skill library can provide. We propose \emph{SkillTTA}, a Test-Time Adaptive Skill Synthesis method that retrieves a small set of training trajectories relevant to the current task and synthesizes them into a temporary, task-specific textual skill. The solver model is kept fixed, so adaptation happens entirely through generated context rather than parameter updates. We evaluate the method on SpreadsheetBench, ALFWorld, and BigCodeBench. Compared with static trajectory-to-skill synthesis using GPT-5.5, task-specific skills improve SpreadsheetBench Pass@1 from 0.397 to 0.505 and BigCodeBench Pass@1 from 0.517 to 0.651. On ALFWorld, the method matches a heavier memory-learning baseline within four points of success rate while producing the shortest successful trajectories among reported methods. Ablations on SpreadsheetBench further show that synthesized skills outperform raw trajectory prompting, that top-$k$ retrieval should stay small, and that failed trajectories are especially useful because they expose recurring evaluator-facing mistakes.
title Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.16986