Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zeyuan, Kim, Jeonghye, Luo, Xufang, Li, Dongsheng, Yang, Yuqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
by: Luo, Xufang, et al.
Published: (2025)
by: Luo, Xufang, et al.
Published: (2025)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
A Large-scale Medical Visual Task Adaptation Benchmark
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
State Contamination in Memory-Augmented LLM Agents
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
by: Wang, Junzhe, et al.
Published: (2026)
by: Wang, Junzhe, et al.
Published: (2026)
Exploring Multi-Modal Data with Tool-Augmented LLM Agents for Precise Causal Discovery
by: Shen, ChengAo, et al.
Published: (2024)
by: Shen, ChengAo, et al.
Published: (2024)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
by: Feng, Xinshun, et al.
Published: (2026)
by: Feng, Xinshun, et al.
Published: (2026)
Group-in-Group Policy Optimization for LLM Agent Training
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
HAEPO: History-Aggregated Exploratory Policy Optimization
by: Trivedi, Gaurish, et al.
Published: (2025)
by: Trivedi, Gaurish, et al.
Published: (2025)
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
by: Yang, Zhaohui, et al.
Published: (2025)
by: Yang, Zhaohui, et al.
Published: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Divergence-Augmented Policy Optimization
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
by: Wang, Hanyong, et al.
Published: (2026)
by: Wang, Hanyong, et al.
Published: (2026)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
by: Kagaya, Tomoyuki, et al.
Published: (2024)
by: Kagaya, Tomoyuki, et al.
Published: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
by: Si, Wenwen, et al.
Published: (2025)
by: Si, Wenwen, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
MLCopilot: Unleashing the Power of Large Language Models in Solving Machine Learning Tasks
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
by: Liu, Zongkai, et al.
Published: (2024)
by: Liu, Zongkai, et al.
Published: (2024)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
Policy Induction: Predicting Startup Success via Explainable Memory-Augmented In-Context Learning
by: Mu, Xianling, et al.
Published: (2025)
by: Mu, Xianling, et al.
Published: (2025)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
by: Wang, Boxiao, et al.
Published: (2026)
by: Wang, Boxiao, et al.
Published: (2026)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
by: Han, Siwei, et al.
Published: (2025)
by: Han, Siwei, et al.
Published: (2025)
Similar Items
-
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
by: Kim, Jeonghye, et al.
Published: (2026) -
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
by: Luo, Xufang, et al.
Published: (2025) -
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
by: Mo, Shentong, et al.
Published: (2026) -
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025) -
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)