VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Guobin, Zhao, Chenxiao, Cheng, Xiang, Huang, Lei, Yu, Xing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Soft Sequence Policy Optimization
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
von: Huang, Luke J., et al.
Veröffentlicht: (2026)
Variational Delayed Policy Optimization
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
Group-in-Group Policy Optimization for LLM Agent Training
von: Feng, Lang, et al.
Veröffentlicht: (2025)
von: Feng, Lang, et al.
Veröffentlicht: (2025)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
Group Sequence Policy Optimization
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
von: Li, Chenliang, et al.
Veröffentlicht: (2025)
von: Li, Chenliang, et al.
Veröffentlicht: (2025)
Reflective Policy Optimization
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
Smooth Gate Functions for Soft Advantage Policy Optimization
von: Denisov, Egor, et al.
Veröffentlicht: (2026)
von: Denisov, Egor, et al.
Veröffentlicht: (2026)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Learning to Reason under Off-Policy Guidance
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yan, Jianhao, et al.
Veröffentlicht: (2025)
Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
von: Wen, Muning, et al.
Veröffentlicht: (2024)
von: Wen, Muning, et al.
Veröffentlicht: (2024)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Clustering Context in Off-Policy Evaluation
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
Zero-Shot Off-Policy Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
Concept-driven Off Policy Evaluation
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
A Unifying View of Coverage in Linear Off-Policy Evaluation
von: Amortila, Philip, et al.
Veröffentlicht: (2026)
von: Amortila, Philip, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025) -
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026) -
Soft Sequence Policy Optimization
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026) -
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)