PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training
Fuente:
arXiv
Saved in:
| Main Authors: | Lv, Mingrui, Liu, Hangzhi, Luo, Zhi, Zhang, Hongjie, Ou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026)
by: Zhou, Huilin, et al.
Published: (2026)
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
by: Huang, Chenyu, et al.
Published: (2026)
by: Huang, Chenyu, et al.
Published: (2026)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
by: Carvalho, Tales H., et al.
Published: (2024)
by: Carvalho, Tales H., et al.
Published: (2024)
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
by: Yu, Yajie, et al.
Published: (2025)
by: Yu, Yajie, et al.
Published: (2025)
DiPRL: Learning Discrete Programmatic Policies via Architecture Entropy Regularization
by: Hu, Chengpeng, et al.
Published: (2026)
by: Hu, Chengpeng, et al.
Published: (2026)
Searching for Programmatic Policies in Semantic Spaces
by: Moraes, Rubens O., et al.
Published: (2024)
by: Moraes, Rubens O., et al.
Published: (2024)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
by: Feng, Xinshun, et al.
Published: (2026)
by: Feng, Xinshun, et al.
Published: (2026)
Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
by: Kohler, Hector, et al.
Published: (2024)
by: Kohler, Hector, et al.
Published: (2024)
Self-Evolving LLMs via Continual Instruction Tuning
by: Kang, Jiazheng, et al.
Published: (2025)
by: Kang, Jiazheng, et al.
Published: (2025)
Evolving Programmatic Skill Networks
by: Shi, Haochen, et al.
Published: (2026)
by: Shi, Haochen, et al.
Published: (2026)
Data-Efficient Training by Evolved Sampling
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
by: Yu, Zhiyin, et al.
Published: (2026)
by: Yu, Zhiyin, et al.
Published: (2026)
Meta-Evolve: Continuous Robot Evolution for One-to-many Policy Transfer
by: Liu, Xingyu, et al.
Published: (2024)
by: Liu, Xingyu, et al.
Published: (2024)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
by: Liu, Max, et al.
Published: (2024)
by: Liu, Max, et al.
Published: (2024)
PolicyBank: Evolving Policy Understanding for LLM Agents
by: Choi, Jihye, et al.
Published: (2026)
by: Choi, Jihye, et al.
Published: (2026)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast
by: Bogdanov, Igor, et al.
Published: (2026)
by: Bogdanov, Igor, et al.
Published: (2026)
Generalized Incremental Learning under Concept Drift across Evolving Data Streams
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
Diving into Self-Evolving Training for Multimodal Reasoning
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
by: Yu, Wenhao, et al.
Published: (2025)
by: Yu, Wenhao, et al.
Published: (2025)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
by: Lu, Yifei, et al.
Published: (2025)
by: Lu, Yifei, et al.
Published: (2025)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Symbolic Learning Enables Self-Evolving Agents
by: Zhou, Wangchunshu, et al.
Published: (2024)
by: Zhou, Wangchunshu, et al.
Published: (2024)
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
by: Zhang, Wenqi, et al.
Published: (2024)
by: Zhang, Wenqi, et al.
Published: (2024)
A Differential Geometric View and Explainability of GNN on Evolving Graphs
by: Liu, Yazheng, et al.
Published: (2024)
by: Liu, Yazheng, et al.
Published: (2024)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
SPELL: Synthesis of Programmatic Edits using LLMs
by: Ramos, Daniel, et al.
Published: (2026)
by: Ramos, Daniel, et al.
Published: (2026)
CMKL: Modality-Aware Continual Learning for Evolving Biomedical Knowledge Graphs
by: Radwan, Yousef A., et al.
Published: (2026)
by: Radwan, Yousef A., et al.
Published: (2026)
GCAL: Adapting Graph Models to Evolving Domain Shifts
by: Qiao, Ziyue, et al.
Published: (2025)
by: Qiao, Ziyue, et al.
Published: (2025)
Policy and World Modeling Co-Training for Language Agents
by: Lu, Ning, et al.
Published: (2026)
by: Lu, Ning, et al.
Published: (2026)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
by: Dionisopoulos, Lucas, et al.
Published: (2026)
by: Dionisopoulos, Lucas, et al.
Published: (2026)
Temporal Generalization Estimation in Evolving Graphs
by: Lu, Bin, et al.
Published: (2024)
by: Lu, Bin, et al.
Published: (2024)
RewardHarness: Self-Evolving Agentic Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
by: Liu, Tennison, et al.
Published: (2025)
by: Liu, Tennison, et al.
Published: (2025)
Similar Items
-
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026) -
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026) -
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
by: Huang, Chenyu, et al.
Published: (2026) -
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026) -
Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces
by: Carvalho, Tales H., et al.
Published: (2024)