Logit Dynamics in Softmax Policy Gradient Methods
Fuente:
arXiv
Saved in:
| Main Author: | Li, Yingru |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Divergence-Augmented Policy Optimization
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
by: Guan, Zhong, et al.
Published: (2026)
by: Guan, Zhong, et al.
Published: (2026)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)
by: Kim, Hoyong, et al.
Published: (2023)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Matrix Low-Rank Approximation For Policy Gradient Methods
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
by: Panangaden, Prakash, et al.
Published: (2023)
by: Panangaden, Prakash, et al.
Published: (2023)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness
by: Xuan, Hao, et al.
Published: (2025)
by: Xuan, Hao, et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
An Effective Dynamic Gradient Calibration Method for Continual Learning
by: Lin, Weichen, et al.
Published: (2024)
by: Lin, Weichen, et al.
Published: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Logit Distance Bounds Representational Similarity
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
by: Chou, Yuhong, et al.
Published: (2024)
by: Chou, Yuhong, et al.
Published: (2024)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Stabilizing Policy Gradient Methods via Reward Profiling
by: Ahmed, Shihab, et al.
Published: (2025)
by: Ahmed, Shihab, et al.
Published: (2025)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Policy Gradient with Kernel Quadrature
by: Hayakawa, Satoshi, et al.
Published: (2023)
by: Hayakawa, Satoshi, et al.
Published: (2023)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment
by: Kitkana, Chayanon, et al.
Published: (2026)
by: Kitkana, Chayanon, et al.
Published: (2026)
Formalising the Logit Shift Induced by LoRA: A Technical Note
by: Shi, Xiang, et al.
Published: (2026)
by: Shi, Xiang, et al.
Published: (2026)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
by: Fan, Chenghao, et al.
Published: (2024)
by: Fan, Chenghao, et al.
Published: (2024)
Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium
by: Lee-Jenkins, Christopher R.
Published: (2025)
by: Lee-Jenkins, Christopher R.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Similar Items
-
Divergence-Augmented Policy Optimization
by: Wang, Qing, et al.
Published: (2025) -
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024) -
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025) -
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
by: Guan, Zhong, et al.
Published: (2026) -
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)