SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Zhi, Gu, Yu, Liu, Wei, Teh, Yee Whye, Lee, Wee Sun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories
by: Kang, Liwei, et al.
Published: (2026)
by: Kang, Liwei, et al.
Published: (2026)
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
Are We Evaluating the Edit Locality of LLM Model Editing Properly?
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)
by: Sims, Anya, et al.
Published: (2024)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
SymDiff: Equivariant Diffusion via Stochastic Symmetrisation
by: Zhang, Leo, et al.
Published: (2024)
by: Zhang, Leo, et al.
Published: (2024)
Beyond Imitation: Reinforcement Learning for Active Latent Planning
by: Zheng, Zhi, et al.
Published: (2026)
by: Zheng, Zhi, et al.
Published: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents
by: Lee, Dongjun, et al.
Published: (2025)
by: Lee, Dongjun, et al.
Published: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
Meta-Learning Objectives for Preference Optimization
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
by: Lai, Yuhang, et al.
Published: (2026)
by: Lai, Yuhang, et al.
Published: (2026)
L3Ms -- Lagrange Large Language Models
by: Dhillon, Guneet S., et al.
Published: (2024)
by: Dhillon, Guneet S., et al.
Published: (2024)
Incorporating Unlabelled Data into Bayesian Neural Networks
by: Sharma, Mrinank, et al.
Published: (2023)
by: Sharma, Mrinank, et al.
Published: (2023)
Is Model Editing Built on Sand? Revealing Its Illusory Success and Fragile Foundation
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick
by: Fu, Jiayi, et al.
Published: (2024)
by: Fu, Jiayi, et al.
Published: (2024)
dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models
by: Wan, Zhengyan, et al.
Published: (2026)
by: Wan, Zhengyan, et al.
Published: (2026)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
by: Galashov, Alexandre, et al.
Published: (2024)
by: Galashov, Alexandre, et al.
Published: (2024)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
Selective Safety Steering via Value-Filtered Decoding
by: Einbinder, Bat-Sheva, et al.
Published: (2026)
by: Einbinder, Bat-Sheva, et al.
Published: (2026)
Manifold Aware Denoising Score Matching (MAD)
by: Levy-Jurgenson, Alona, et al.
Published: (2026)
by: Levy-Jurgenson, Alona, et al.
Published: (2026)
Joint Beamforming and Integer User Association using a GNN with Gumbel-Softmax Reparameterizations
by: Lyu, Qing, et al.
Published: (2025)
by: Lyu, Qing, et al.
Published: (2025)
Deep Thinking by Markov Chain of Continuous Thoughts
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
Rao-Blackwellised Reparameterisation Gradients
by: Lam, Kevin H., et al.
Published: (2025)
by: Lam, Kevin H., et al.
Published: (2025)
Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
by: Liu, Qiyuan, et al.
Published: (2025)
by: Liu, Qiyuan, et al.
Published: (2025)
Reasoning-CV: Fine-tuning Powerful Reasoning LLMs for Knowledge-Assisted Claim Verification
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
Amortized Probabilistic Detection of Communities in Graphs
by: Wang, Yueqi, et al.
Published: (2020)
by: Wang, Yueqi, et al.
Published: (2020)
EvIL: Evolution Strategies for Generalisable Imitation Learning
by: Sapora, Silvia, et al.
Published: (2024)
by: Sapora, Silvia, et al.
Published: (2024)
A Reparameterized Discrete Diffusion Model for Text Generation
by: Zheng, Lin, et al.
Published: (2023)
by: Zheng, Lin, et al.
Published: (2023)
Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts
by: Chen, Shengzhuang, et al.
Published: (2024)
by: Chen, Shengzhuang, et al.
Published: (2024)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
SigmaDock: Untwisting Molecular Docking With Fragment-Based SE(3) Diffusion
by: Prat, Alvaro, et al.
Published: (2025)
by: Prat, Alvaro, et al.
Published: (2025)
Meta Flow Maps enable scalable reward alignment
by: Potaptchik, Peter, et al.
Published: (2026)
by: Potaptchik, Peter, et al.
Published: (2026)
Prompting Strategies for Enabling Large Language Models to Infer Causation from Correlation
by: Sgouritsa, Eleni, et al.
Published: (2024)
by: Sgouritsa, Eleni, et al.
Published: (2024)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
Similar Items
-
LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories
by: Kang, Liwei, et al.
Published: (2026) -
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
by: Liu, Wei, et al.
Published: (2026) -
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
by: Nguyen-Hien, T. Duy, et al.
Published: (2025) -
Are We Evaluating the Edit Locality of LLM Model Editing Properly?
by: Liu, Wei, et al.
Published: (2026) -
The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning
by: Sims, Anya, et al.
Published: (2024)