PolicyLong: Towards On-Policy Context Extension
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Junlong, Chen, Ziyang, Wu, Xing, Gao, Chaochen, Yu, TingHao, Zhang, Feng, Hu, Songlin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
NExtLong: Toward Effective Long-Context Training without Long Documents
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
by: Gao, Chaochen, et al.
Published: (2024)
by: Gao, Chaochen, et al.
Published: (2024)
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
by: Lu, Dekun, et al.
Published: (2025)
by: Lu, Dekun, et al.
Published: (2025)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
by: Yang, Ningyuan, et al.
Published: (2025)
by: Yang, Ningyuan, et al.
Published: (2025)
Libra: Large Chinese-based Safeguard for AI Content
by: Chen, Ziyang, et al.
Published: (2025)
by: Chen, Ziyang, et al.
Published: (2025)
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
by: Chen, Xianwei, et al.
Published: (2026)
by: Chen, Xianwei, et al.
Published: (2026)
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026)
by: Yu, Tianrun, et al.
Published: (2026)
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
by: He, Shuo, et al.
Published: (2026)
by: He, Shuo, et al.
Published: (2026)
eDOC: Explainable Decoding Out-of-domain Cell Types with Evidential Learning
by: Wu, Chaochen, et al.
Published: (2024)
by: Wu, Chaochen, et al.
Published: (2024)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
by: Yu, Yajie, et al.
Published: (2025)
by: Yu, Yajie, et al.
Published: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2025)
by: Gao, Chen-Xiao, et al.
Published: (2025)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
Segment-Aligned Policy Optimization for Multi-Modal Reasoning
by: Gao, Lei, et al.
Published: (2026)
by: Gao, Lei, et al.
Published: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
by: Jia, Chengxing, et al.
Published: (2024)
by: Jia, Chengxing, et al.
Published: (2024)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
by: Ye, Chenlu, et al.
Published: (2026)
by: Ye, Chenlu, et al.
Published: (2026)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
by: Nam, Hyunji, et al.
Published: (2024)
by: Nam, Hyunji, et al.
Published: (2024)
Random Policy Enables In-Context Reinforcement Learning within Trust Horizons
by: Chen, Weiqin, et al.
Published: (2024)
by: Chen, Weiqin, et al.
Published: (2024)
Neural Paging: Learning Context Management Policies for Turing-Complete Agents
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Policy Constraint by Only Support Constraint for Offline Reinforcement Learning
by: Gao, Yunkai, et al.
Published: (2025)
by: Gao, Yunkai, et al.
Published: (2025)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
by: Hu, Hao, et al.
Published: (2025)
by: Hu, Hao, et al.
Published: (2025)
A Causal Lens for Learning Long-term Fair Policies
by: Lear, Jacob, et al.
Published: (2025)
by: Lear, Jacob, et al.
Published: (2025)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
by: Jia, Nan, et al.
Published: (2026)
by: Jia, Nan, et al.
Published: (2026)
Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
by: Yao, Yongqiang, et al.
Published: (2025)
by: Yao, Yongqiang, et al.
Published: (2025)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
LLoCO: Learning Long Contexts Offline
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
Program Machine Policy: Addressing Long-Horizon Tasks by Integrating Program Synthesis and State Machines
by: Lin, Yu-An, et al.
Published: (2023)
by: Lin, Yu-An, et al.
Published: (2023)
Similar Items
-
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026) -
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
by: Jia, Junlong, et al.
Published: (2025) -
NExtLong: Toward Effective Long-Context Training without Long Documents
by: Gao, Chaochen, et al.
Published: (2025) -
Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
by: Gao, Chaochen, et al.
Published: (2024) -
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
by: Gao, Chaochen, et al.
Published: (2025)