Soft Sequence Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Glazyrina, Svetlana, Kryzhanovskiy, Maksim, Ischenko, Roman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Smooth Gate Functions for Soft Advantage Policy Optimization
von: Denisov, Egor, et al.
Veröffentlicht: (2026)
von: Denisov, Egor, et al.
Veröffentlicht: (2026)
Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
von: Kryzhanovskiy, Maksim, et al.
Veröffentlicht: (2025)
von: Kryzhanovskiy, Maksim, et al.
Veröffentlicht: (2025)
Topic Modelling Black Box Optimization
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Soft Adaptive Policy Optimization
von: Gao, Chang, et al.
Veröffentlicht: (2025)
von: Gao, Chang, et al.
Veröffentlicht: (2025)
Group Sequence Policy Optimization
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
Zero-Shot Off-Policy Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
Direct Soft-Policy Sampling via Langevin Dynamics
von: Ki, Donghyeon, et al.
Veröffentlicht: (2026)
von: Ki, Donghyeon, et al.
Veröffentlicht: (2026)
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
Soft Graph Clustering for single-cell RNA Sequencing Data
von: Xu, Ping, et al.
Veröffentlicht: (2025)
von: Xu, Ping, et al.
Veröffentlicht: (2025)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
von: Kang, Hyeongyu, et al.
Veröffentlicht: (2025)
von: Kang, Hyeongyu, et al.
Veröffentlicht: (2025)
Wasserstein Policy Optimization
von: Pfau, David, et al.
Veröffentlicht: (2025)
von: Pfau, David, et al.
Veröffentlicht: (2025)
Reflective Policy Optimization
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Soft Preference Optimization: Aligning Language Models to Expert Distributions
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
von: Szarvas, Bálint László, et al.
Veröffentlicht: (2025)
von: Szarvas, Bálint László, et al.
Veröffentlicht: (2025)
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
Single-stream Policy Optimization
von: Xu, Zhongwen, et al.
Veröffentlicht: (2025)
von: Xu, Zhongwen, et al.
Veröffentlicht: (2025)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
Variational Delayed Policy Optimization
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
Divergence-Augmented Policy Optimization
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Decision Flow Policy Optimization
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
Fractal Landscapes in Policy Optimization
von: Wang, Tao, et al.
Veröffentlicht: (2023)
von: Wang, Tao, et al.
Veröffentlicht: (2023)
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025)
von: Sane, Soham
Veröffentlicht: (2025)
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
von: Su, Zelal, et al.
Veröffentlicht: (2026)
von: Su, Zelal, et al.
Veröffentlicht: (2026)
Local Entropy Search over Descent Sequences for Bayesian Optimization
von: Stenger, David, et al.
Veröffentlicht: (2025)
von: Stenger, David, et al.
Veröffentlicht: (2025)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
Gradient Extrapolation-Based Policy Optimization
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
Complexity-Regularized Proximal Policy Optimization
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Smooth Gate Functions for Soft Advantage Policy Optimization
von: Denisov, Egor, et al.
Veröffentlicht: (2026) -
Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
von: Kryzhanovskiy, Maksim, et al.
Veröffentlicht: (2025) -
Topic Modelling Black Box Optimization
von: Akramov, Roman, et al.
Veröffentlicht: (2025) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025) -
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)