Learning Agentic Policy from Action Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Yuxiang, Wang, Zengbin, Wang, Yong, Yang, Shidong, Ma, Ziyu, Chen, Guanhua, Sun, Zonghua, Wu, Liaoni, Chu, Xiangxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
by: Ji, Yuxiang, et al.
Published: (2026)
by: Ji, Yuxiang, et al.
Published: (2026)
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
by: Ma, Ziyu, et al.
Published: (2026)
by: Ma, Ziyu, et al.
Published: (2026)
Tree Search for LLM Agent Reinforcement Learning
by: Ji, Yuxiang, et al.
Published: (2025)
by: Ji, Yuxiang, et al.
Published: (2025)
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
by: Yang, Shidong, et al.
Published: (2026)
by: Yang, Shidong, et al.
Published: (2026)
AR-MAP: Are Autoregressive Large Language Models Implicit Teachers for Diffusion Large Language Models?
by: Lin, Liang, et al.
Published: (2026)
by: Lin, Liang, et al.
Published: (2026)
Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution
by: Xiong, Feng, et al.
Published: (2026)
by: Xiong, Feng, et al.
Published: (2026)
GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation
by: Dai, Yanqi, et al.
Published: (2026)
by: Dai, Yanqi, et al.
Published: (2026)
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
by: Zhang, Ru, et al.
Published: (2026)
by: Zhang, Ru, et al.
Published: (2026)
SSL: Sweet Spot Learning for Differentiated Guidance in Agentic Optimization
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
by: Wang, Zengbin, et al.
Published: (2026)
by: Wang, Zengbin, et al.
Published: (2026)
Pi-SQL: Enhancing Text-to-SQL with Fine-Grained Guidance from Pivot Programming Languages
by: chi, Yongdong, et al.
Published: (2025)
by: chi, Yongdong, et al.
Published: (2025)
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
by: Xiao, Zeguan, et al.
Published: (2026)
by: Xiao, Zeguan, et al.
Published: (2026)
Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation
by: Shi, Dingyuan, et al.
Published: (2024)
by: Shi, Dingyuan, et al.
Published: (2024)
Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability
by: He, Boyong, et al.
Published: (2025)
by: He, Boyong, et al.
Published: (2025)
Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector
by: He, Boyong, et al.
Published: (2025)
by: He, Boyong, et al.
Published: (2025)
Game4Loc: A UAV Geo-Localization Benchmark from Game Data
by: Ji, Yuxiang, et al.
Published: (2024)
by: Ji, Yuxiang, et al.
Published: (2024)
Breaking Block Boundaries: Anchor-based History-stable Decoding for Diffusion Large Language Models
by: Zou, Shun, et al.
Published: (2026)
by: Zou, Shun, et al.
Published: (2026)
VisLanding: Monocular 3D Perception for UAV Safe Landing via Depth-Normal Synergy
by: Tan, Zhuoyue, et al.
Published: (2025)
by: Tan, Zhuoyue, et al.
Published: (2025)
Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
by: Wang, Daoyu, et al.
Published: (2026)
by: Wang, Daoyu, et al.
Published: (2026)
AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting
by: Li, Renda, et al.
Published: (2025)
by: Li, Renda, et al.
Published: (2025)
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
by: Dou, Zhihao, et al.
Published: (2025)
by: Dou, Zhihao, et al.
Published: (2025)
KnowRA: Knowledge Retrieval Augmented Method for Document-level Relation Extraction with Comprehensive Reasoning Abilities
by: Mai, Chengcheng, et al.
Published: (2024)
by: Mai, Chengcheng, et al.
Published: (2024)
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design
by: Chen, Shengchao, et al.
Published: (2026)
by: Chen, Shengchao, et al.
Published: (2026)
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
by: Xiao, Zeguan, et al.
Published: (2026)
by: Xiao, Zeguan, et al.
Published: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
by: Xiong, Feng, et al.
Published: (2025)
by: Xiong, Feng, et al.
Published: (2025)
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
by: Chen, Dingwei, et al.
Published: (2026)
by: Chen, Dingwei, et al.
Published: (2026)
Policy Learning with a Natural Language Action Space: A Causal Approach
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
SWAG: Storytelling With Action Guidance
by: Patel, Zeeshan, et al.
Published: (2024)
by: Patel, Zeeshan, et al.
Published: (2024)
Demystifying Reinforcement Learning in Agentic Reasoning
by: Yu, Zhaochen, et al.
Published: (2025)
by: Yu, Zhaochen, et al.
Published: (2025)
Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detection
by: He, Boyong, et al.
Published: (2025)
by: He, Boyong, et al.
Published: (2025)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Similar Items
-
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
by: Ji, Yuxiang, et al.
Published: (2026) -
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
by: Ma, Ziyu, et al.
Published: (2026) -
Tree Search for LLM Agent Reinforcement Learning
by: Ji, Yuxiang, et al.
Published: (2025) -
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
by: Yang, Shidong, et al.
Published: (2026) -
AR-MAP: Are Autoregressive Large Language Models Implicit Teachers for Diffusion Large Language Models?
by: Lin, Liang, et al.
Published: (2026)