Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Huichi, Chen, Yihang, Guo, Siyuan, Yan, Xue, Lee, Kin Hei, Wang, Zihan, Lee, Ka Yiu, Zhang, Guchun, Shao, Kun, Yang, Linyi, Wang, Jun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912552754085888
author Zhou, Huichi
Chen, Yihang
Guo, Siyuan
Yan, Xue
Lee, Kin Hei
Wang, Zihan
Lee, Ka Yiu
Zhang, Guchun
Shao, Kun
Yang, Linyi
Wang, Jun
author_facet Zhou, Huichi
Chen, Yihang
Guo, Siyuan
Yan, Xue
Lee, Kin Hei
Wang, Zihan
Lee, Ka Yiu
Zhang, Guchun
Shao, Kun
Yang, Linyi
Wang, Jun
contents In this paper, we introduce a novel learning paradigm for Adaptive Large Language Model (LLM) agents that eliminates the need for fine-tuning the underlying LLMs. Existing approaches are often either rigid, relying on static, handcrafted reflection workflows, or computationally intensive, requiring gradient updates of LLM model parameters. In contrast, our method enables low-cost continual adaptation via memory-based online reinforcement learning. We formalise this as a Memory-augmented Markov Decision Process (M-MDP), equipped with a neural case-selection policy to guide action decisions. Past experiences are stored in an episodic memory, either differentiable or non-parametric. The policy is continually updated based on environmental feedback through a memory rewriting mechanism, whereas policy improvement is achieved through efficient memory reading (retrieval). We instantiate our agent model in the deep research setting, namely \emph{Memento}, which attains top-1 on GAIA validation ($87.88\%$ Pass@$3$) and $79.40\%$ on the test set. It reaches $66.6\%$ F1 and $80.4\%$ PM on the DeepResearcher dataset, outperforming the state-of-the-art training-based method, while case-based memory adds $4.7\%$ to $9.6\%$ absolute points on out-of-distribution tasks. Our approach offers a scalable and efficient pathway for developing generalist LLM agents capable of continuous, real-time learning without gradient updates, advancing machine learning towards open-ended skill acquisition and deep research scenarios. The code is available at https://github.com/Agent-on-the-Fly/Memento.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16153
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Zhou, Huichi
Chen, Yihang
Guo, Siyuan
Yan, Xue
Lee, Kin Hei
Wang, Zihan
Lee, Ka Yiu
Zhang, Guchun
Shao, Kun
Yang, Linyi
Wang, Jun
Machine Learning
Computation and Language
In this paper, we introduce a novel learning paradigm for Adaptive Large Language Model (LLM) agents that eliminates the need for fine-tuning the underlying LLMs. Existing approaches are often either rigid, relying on static, handcrafted reflection workflows, or computationally intensive, requiring gradient updates of LLM model parameters. In contrast, our method enables low-cost continual adaptation via memory-based online reinforcement learning. We formalise this as a Memory-augmented Markov Decision Process (M-MDP), equipped with a neural case-selection policy to guide action decisions. Past experiences are stored in an episodic memory, either differentiable or non-parametric. The policy is continually updated based on environmental feedback through a memory rewriting mechanism, whereas policy improvement is achieved through efficient memory reading (retrieval). We instantiate our agent model in the deep research setting, namely \emph{Memento}, which attains top-1 on GAIA validation ($87.88\%$ Pass@$3$) and $79.40\%$ on the test set. It reaches $66.6\%$ F1 and $80.4\%$ PM on the DeepResearcher dataset, outperforming the state-of-the-art training-based method, while case-based memory adds $4.7\%$ to $9.6\%$ absolute points on out-of-distribution tasks. Our approach offers a scalable and efficient pathway for developing generalist LLM agents capable of continuous, real-time learning without gradient updates, advancing machine learning towards open-ended skill acquisition and deep research scenarios. The code is available at https://github.com/Agent-on-the-Fly/Memento.
title Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2508.16153