APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yibo, Yang, Jiashuo, Zheng, Zhi, Hu, Zhiyuan, Sui, Yuan, Wang, Shizun, He, Yufei, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
by: Sui, Yuan, et al.
Published: (2026)
by: Sui, Yuan, et al.
Published: (2026)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary Learning
by: He, Yufei, et al.
Published: (2026)
by: He, Yufei, et al.
Published: (2026)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
by: Lou, Zhanzhi, et al.
Published: (2026)
by: Lou, Zhanzhi, et al.
Published: (2026)
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
by: Sui, Yuan, et al.
Published: (2025)
by: Sui, Yuan, et al.
Published: (2025)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering
by: Sui, Yuan, et al.
Published: (2024)
by: Sui, Yuan, et al.
Published: (2024)
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
by: Sui, Yuan, et al.
Published: (2026)
by: Sui, Yuan, et al.
Published: (2026)
Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
APEX-Agents
by: Vidgen, Bertie, et al.
Published: (2026)
by: Vidgen, Bertie, et al.
Published: (2026)
HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
by: Wang, Haochen, et al.
Published: (2026)
by: Wang, Haochen, et al.
Published: (2026)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
by: Chen, Hui, et al.
Published: (2025)
by: Chen, Hui, et al.
Published: (2025)
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question Answering
by: Sui, Yuan, et al.
Published: (2024)
by: Sui, Yuan, et al.
Published: (2024)
APEX: Agent Payment Execution with Policy for Autonomous Agent API Access
by: Uddin, Mohd Safwan, et al.
Published: (2026)
by: Uddin, Mohd Safwan, et al.
Published: (2026)
KLong: Training LLM Agent for Extremely Long-horizon Tasks
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
by: Guo, Qianyun, et al.
Published: (2026)
by: Guo, Qianyun, et al.
Published: (2026)
LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts
by: Zeng, Yibo, et al.
Published: (2024)
by: Zeng, Yibo, et al.
Published: (2024)
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs
by: He, Yufei, et al.
Published: (2024)
by: He, Yufei, et al.
Published: (2024)
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
by: Yu, Yajie, et al.
Published: (2025)
by: Yu, Yajie, et al.
Published: (2025)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
by: Feng, Xinshun, et al.
Published: (2026)
by: Feng, Xinshun, et al.
Published: (2026)
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training
by: Lv, Mingrui, et al.
Published: (2025)
by: Lv, Mingrui, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
by: Liu, Tennison, et al.
Published: (2025)
by: Liu, Tennison, et al.
Published: (2025)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
by: Wu, Jiayun, et al.
Published: (2025)
by: Wu, Jiayun, et al.
Published: (2025)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
by: Zhou, Huilin, et al.
Published: (2026)
by: Zhou, Huilin, et al.
Published: (2026)
Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
by: Zhang, Jintian, et al.
Published: (2023)
by: Zhang, Jintian, et al.
Published: (2023)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification
by: Hu, Jun, et al.
Published: (2025)
by: Hu, Jun, et al.
Published: (2025)
APEX: Probing Neural Networks via Activation Perturbation
by: Ren, Tao, et al.
Published: (2026)
by: Ren, Tao, et al.
Published: (2026)
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
by: Fan, Zhiyuan, et al.
Published: (2026)
by: Fan, Zhiyuan, et al.
Published: (2026)
Similar Items
-
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
by: Sui, Yuan, et al.
Published: (2026) -
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026) -
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary Learning
by: He, Yufei, et al.
Published: (2026) -
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
by: Li, Yibo, et al.
Published: (2026) -
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
by: Lou, Zhanzhi, et al.
Published: (2026)