Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Shuo, Feng, Lang, Wei, Qi, Cheng, Xin, Feng, Lei, An, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
von: He, Shuo, et al.
Veröffentlicht: (2026)
von: He, Shuo, et al.
Veröffentlicht: (2026)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
von: Feng, Lang, et al.
Veröffentlicht: (2026)
von: Feng, Lang, et al.
Veröffentlicht: (2026)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
von: Zhao, Yujie, et al.
Veröffentlicht: (2026)
von: Zhao, Yujie, et al.
Veröffentlicht: (2026)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
von: He, Zelin, et al.
Veröffentlicht: (2026)
von: He, Zelin, et al.
Veröffentlicht: (2026)
Program Machine Policy: Addressing Long-Horizon Tasks by Integrating Program Synthesis and State Machines
von: Lin, Yu-An, et al.
Veröffentlicht: (2023)
von: Lin, Yu-An, et al.
Veröffentlicht: (2023)
Double Horizon Model-Based Policy Optimization
von: Kubo, Akihiro, et al.
Veröffentlicht: (2025)
von: Kubo, Akihiro, et al.
Veröffentlicht: (2025)
SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks
von: Wen, Yongyan, et al.
Veröffentlicht: (2024)
von: Wen, Yongyan, et al.
Veröffentlicht: (2024)
AgentOCR: Reimagining Agent History via Optical Self-Compression
von: Feng, Lang, et al.
Veröffentlicht: (2026)
von: Feng, Lang, et al.
Veröffentlicht: (2026)
From Semantics to Hierarchy: A Hybrid Euclidean-Tangent-Hyperbolic Space Model for Temporal Knowledge Graph Reasoning
von: Feng, Siling, et al.
Veröffentlicht: (2024)
von: Feng, Siling, et al.
Veröffentlicht: (2024)
SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2026)
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2026)
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
von: Han, Xinchen, et al.
Veröffentlicht: (2026)
von: Han, Xinchen, et al.
Veröffentlicht: (2026)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement Learning
von: Xue, Zhenghai, et al.
Veröffentlicht: (2025)
von: Xue, Zhenghai, et al.
Veröffentlicht: (2025)
LHAW: Controllable Underspecification for Long-Horizon Tasks
von: Pu, George, et al.
Veröffentlicht: (2026)
von: Pu, George, et al.
Veröffentlicht: (2026)
Learning for Long-Horizon Planning via Neuro-Symbolic Abductive Imitation
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2026)
PepEVOLVE: Position-Aware Dynamic Peptide Optimization via Group-Relative Advantage
von: Nguyen, Trieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Trieu, et al.
Veröffentlicht: (2025)
Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation Tree
von: Feng, Lang, et al.
Veröffentlicht: (2024)
von: Feng, Lang, et al.
Veröffentlicht: (2024)
Agentic Reinforced Policy Optimization
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents
von: Pang, Cong, et al.
Veröffentlicht: (2026)
von: Pang, Cong, et al.
Veröffentlicht: (2026)
On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
von: Zheng, Hua, et al.
Veröffentlicht: (2021)
von: Zheng, Hua, et al.
Veröffentlicht: (2021)
PolicyLong: Towards On-Policy Context Extension
von: Jia, Junlong, et al.
Veröffentlicht: (2026)
von: Jia, Junlong, et al.
Veröffentlicht: (2026)
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
von: Hu, Miaobo, et al.
Veröffentlicht: (2026)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
von: Choi, Jinwoo, et al.
Veröffentlicht: (2026)
von: Choi, Jinwoo, et al.
Veröffentlicht: (2026)
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Hindsight Credit Assignment for Long-Horizon LLM Agents
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
von: Tan, Hui-Ze, et al.
Veröffentlicht: (2026)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
Flattening Hierarchies with Policy Bootstrapping
von: Zhou, John L., et al.
Veröffentlicht: (2025)
von: Zhou, John L., et al.
Veröffentlicht: (2025)
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
von: He, Jinmin, et al.
Veröffentlicht: (2025)
von: He, Jinmin, et al.
Veröffentlicht: (2025)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
von: Lei, Shiye, et al.
Veröffentlicht: (2026)
von: Lei, Shiye, et al.
Veröffentlicht: (2026)
Accelerating Task Generalisation with Multi-Level Skill Hierarchies
von: Cannon, Thomas P, et al.
Veröffentlicht: (2024)
von: Cannon, Thomas P, et al.
Veröffentlicht: (2024)
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
von: Xu, Pei, et al.
Veröffentlicht: (2025)
von: Xu, Pei, et al.
Veröffentlicht: (2025)
Autoregressive Policy Optimization for Constrained Allocation Tasks
von: Winkel, David, et al.
Veröffentlicht: (2024)
von: Winkel, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025) -
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
von: He, Shuo, et al.
Veröffentlicht: (2026) -
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
von: Feng, Lang, et al.
Veröffentlicht: (2026) -
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026) -
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
von: Zhao, Yujie, et al.
Veröffentlicht: (2026)