Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Pengcheng, Lin, Jiacheng, Shi, Zhiyi, Wang, Zifeng, He, Luxi, Wu, Yichen, Zhong, Ming, Song, Peiyang, Zhang, Qizheng, Wang, Heng, Xu, Xueqiang, Xu, Hanwen, Han, Pengrui, Zhang, Dylan, Sun, Jiashuo, Yang, Chaoqi, Qian, Kun, Wang, Tian, Hu, Changran, Li, Manling, Li, Quanzheng, Peng, Hao, Wang, Sheng, Shang, Jingbo, Zhang, Chao, You, Jiaxuan, Liu, Liyuan, Lu, Pan, Zhang, Yu, Ji, Heng, Choi, Yejin, Song, Dawn, Sun, Jimeng, Han, Jiawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917323976212480
author Jiang, Pengcheng
Lin, Jiacheng
Shi, Zhiyi
Wang, Zifeng
He, Luxi
Wu, Yichen
Zhong, Ming
Song, Peiyang
Zhang, Qizheng
Wang, Heng
Xu, Xueqiang
Xu, Hanwen
Han, Pengrui
Zhang, Dylan
Sun, Jiashuo
Yang, Chaoqi
Qian, Kun
Wang, Tian
Hu, Changran
Li, Manling
Li, Quanzheng
Peng, Hao
Wang, Sheng
Shang, Jingbo
Zhang, Chao
You, Jiaxuan
Liu, Liyuan
Lu, Pan
Zhang, Yu
Ji, Heng
Choi, Yejin
Song, Dawn
Sun, Jimeng
Han, Jiawei
author_facet Jiang, Pengcheng
Lin, Jiacheng
Shi, Zhiyi
Wang, Zifeng
He, Luxi
Wu, Yichen
Zhong, Ming
Song, Peiyang
Zhang, Qizheng
Wang, Heng
Xu, Xueqiang
Xu, Hanwen
Han, Pengrui
Zhang, Dylan
Sun, Jiashuo
Yang, Chaoqi
Qian, Kun
Wang, Tian
Hu, Changran
Li, Manling
Li, Quanzheng
Peng, Hao
Wang, Sheng
Shang, Jingbo
Zhang, Chao
You, Jiaxuan
Liu, Liyuan
Lu, Pan
Zhang, Yu
Ji, Heng
Choi, Yejin
Song, Dawn
Sun, Jimeng
Han, Jiawei
contents Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
Jiang, Pengcheng
Lin, Jiacheng
Shi, Zhiyi
Wang, Zifeng
He, Luxi
Wu, Yichen
Zhong, Ming
Song, Peiyang
Zhang, Qizheng
Wang, Heng
Xu, Xueqiang
Xu, Hanwen
Han, Pengrui
Zhang, Dylan
Sun, Jiashuo
Yang, Chaoqi
Qian, Kun
Wang, Tian
Hu, Changran
Li, Manling
Li, Quanzheng
Peng, Hao
Wang, Sheng
Shang, Jingbo
Zhang, Chao
You, Jiaxuan
Liu, Liyuan
Lu, Pan
Zhang, Yu
Ji, Heng
Choi, Yejin
Song, Dawn
Sun, Jimeng
Han, Jiawei
Artificial Intelligence
Computation and Language
Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.
title Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.16301