Near-Future Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Chuanyu, Yang, Chenxu, Si, Qingyi, Gu, Naibin, Yao, Dingyu, Lin, Zheng, Fu, Peng, Duan, Nan, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-Evolving Policy Distillation
by: Gu, Naibin, et al.
Published: (2026)
by: Gu, Naibin, et al.
Published: (2026)
EasyVideoR1: Easier RL for Video Understanding
by: Qin, Chuanyu, et al.
Published: (2026)
by: Qin, Chuanyu, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
System 1&2 Synergy via Dynamic Model Interpolation
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
Online Self-Calibration Against Hallucination in Vision-Language Models
by: Chen, Minghui, et al.
Published: (2026)
by: Chen, Minghui, et al.
Published: (2026)
Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models
by: Liu, Xiyu, et al.
Published: (2026)
by: Liu, Xiyu, et al.
Published: (2026)
Are Large Language Models Table-based Fact-Checkers?
by: Zhang, Hanwen, et al.
Published: (2024)
by: Zhang, Hanwen, et al.
Published: (2024)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
Test-time Prompt Intervention
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
Stable Reinforcement Learning for Efficient Reasoning
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
by: Ma, Chiyu, et al.
Published: (2026)
by: Ma, Chiyu, et al.
Published: (2026)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
by: Wang, Zhengcheng, et al.
Published: (2025)
by: Wang, Zhengcheng, et al.
Published: (2025)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
by: Lin, Hongqiang, et al.
Published: (2026)
by: Lin, Hongqiang, et al.
Published: (2026)
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
by: Gao, Ting, et al.
Published: (2026)
by: Gao, Ting, et al.
Published: (2026)
TTAQ: Towards Stable Post-training Quantization in Continuous Domain Adaptation
by: Xiao, Junrui, et al.
Published: (2024)
by: Xiao, Junrui, et al.
Published: (2024)
Near-Optimal Policy Optimization for Correlated Equilibrium in General-Sum Markov Games
by: Cai, Yang, et al.
Published: (2024)
by: Cai, Yang, et al.
Published: (2024)
Task-Centric Policy Optimization from Misaligned Motion Priors
by: Zheng, Ziang, et al.
Published: (2026)
by: Zheng, Ziang, et al.
Published: (2026)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization
by: Qiu, Junbin, et al.
Published: (2025)
by: Qiu, Junbin, et al.
Published: (2025)
RepQuant: Towards Accurate Post-Training Quantization of Large Transformer Models via Scale Reparameterization
by: Li, Zhikai, et al.
Published: (2024)
by: Li, Zhikai, et al.
Published: (2024)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
Continuous Optimization for Feature Selection with Permutation-Invariant Embedding and Policy-Guided Search
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
A Closer Look into LLMs for Table Understanding
by: Wang, Jia, et al.
Published: (2026)
by: Wang, Jia, et al.
Published: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025)
by: Wu, Wenbo, et al.
Published: (2025)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Large Language Models as Generalist Policies for Network Optimization
by: Wu, Duo, et al.
Published: (2025)
by: Wu, Duo, et al.
Published: (2025)
Distributionally Robust Optimization via Generative Ambiguity Modeling
by: Wen, Jiaqi, et al.
Published: (2026)
by: Wen, Jiaqi, et al.
Published: (2026)
Distributionally Robust Optimization via Diffusion Ambiguity Modeling
by: Wen, Jiaqi, et al.
Published: (2025)
by: Wen, Jiaqi, et al.
Published: (2025)
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Dichotomous Diffusion Policy Optimization
by: Liang, Ruiming, et al.
Published: (2025)
by: Liang, Ruiming, et al.
Published: (2025)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
A Regularization-Sharpness Tradeoff for Linear Interpolators
by: Hu, Qingyi, et al.
Published: (2026)
by: Hu, Qingyi, et al.
Published: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Dual Approximation Policy Optimization
by: Xiong, Zhihan, et al.
Published: (2024)
by: Xiong, Zhihan, et al.
Published: (2024)
OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
Similar Items
-
Co-Evolving Policy Distillation
by: Gu, Naibin, et al.
Published: (2026) -
EasyVideoR1: Easier RL for Video Understanding
by: Qin, Chuanyu, et al.
Published: (2026) -
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026) -
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025) -
System 1&2 Synergy via Dynamic Model Interpolation
by: Yang, Chenxu, et al.
Published: (2026)