Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qifan, Ma, Dongyang, Fang, Tianqing, Li, Jia, Tang, Jing, Chen, Nuo, Mi, Haitao, Wang, Yan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913047365287936
author Zhang, Qifan
Ma, Dongyang
Fang, Tianqing
Li, Jia
Tang, Jing
Chen, Nuo
Mi, Haitao
Wang, Yan
author_facet Zhang, Qifan
Ma, Dongyang
Fang, Tianqing
Li, Jia
Tang, Jing
Chen, Nuo
Mi, Haitao
Wang, Yan
contents Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work, we train agents to possess an intrinsic meta-evolution capability to spontaneously learn about unseen environments prior to task execution. To instill this ability, we design an outcome-based reward mechanism that measures how much an agent's self-generated world knowledge improves its success rate on downstream tasks. This reward signal is used exclusively during the training phase to teach the model how to explore and summarize effectively. At inference time, the agent requires no external rewards or human instructions. It spontaneously performs native self-evolution to adapt to unknown environments using its internal parameters. When applied to Qwen3-30B and Seed-OSS-36B, this shift to native evolution yields a 20% performance increase on WebVoyager and WebWalker. Most strikingly, the generated world knowledge even enables a compact 14B Qwen3 model to outperform the unassisted Gemini-2.5-Flash, establishing a new paradigm for truly evolving agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18131
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
Zhang, Qifan
Ma, Dongyang
Fang, Tianqing
Li, Jia
Tang, Jing
Chen, Nuo
Mi, Haitao
Wang, Yan
Artificial Intelligence
Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work, we train agents to possess an intrinsic meta-evolution capability to spontaneously learn about unseen environments prior to task execution. To instill this ability, we design an outcome-based reward mechanism that measures how much an agent's self-generated world knowledge improves its success rate on downstream tasks. This reward signal is used exclusively during the training phase to teach the model how to explore and summarize effectively. At inference time, the agent requires no external rewards or human instructions. It spontaneously performs native self-evolution to adapt to unknown environments using its internal parameters. When applied to Qwen3-30B and Seed-OSS-36B, this shift to native evolution yields a 20% performance increase on WebVoyager and WebWalker. Most strikingly, the generated world knowledge even enables a compact 14B Qwen3 model to outperform the unassisted Gemini-2.5-Flash, establishing a new paradigm for truly evolving agents.
title Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
topic Artificial Intelligence
url https://arxiv.org/abs/2604.18131