Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
Fuente:
arXiv
Saved in:
| Main Author: | Sun, Mingfei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
by: Abdi, Hossein, et al.
Published: (2025)
by: Abdi, Hossein, et al.
Published: (2025)
Gradient Regularized Natural Gradients
by: Dash, Satya Prakash, et al.
Published: (2026)
by: Dash, Satya Prakash, et al.
Published: (2026)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
by: Huo, Yingxiao, et al.
Published: (2026)
by: Huo, Yingxiao, et al.
Published: (2026)
GRADE: Replacing Policy Gradients with Backpropagation for LLM Alignment
by: Nel, Lukas Abrie
Published: (2025)
by: Nel, Lukas Abrie
Published: (2025)
Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
by: Flügel, Katharina, et al.
Published: (2024)
by: Flügel, Katharina, et al.
Published: (2024)
Neural Network Verification with PyRAT
by: Lemesle, Augustin, et al.
Published: (2024)
by: Lemesle, Augustin, et al.
Published: (2024)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Generalized Euler Logarithm and its Applications in Machine Learning: Natural Gradient, Backpropagation, Generalized EG, Mirror Descent and OLPS
by: Cichocki, Andrzej
Published: (2025)
by: Cichocki, Andrzej
Published: (2025)
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
by: Yang, Ningyuan, et al.
Published: (2025)
by: Yang, Ningyuan, et al.
Published: (2025)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
Practical Boolean Backpropagation
by: Golbert, Simon
Published: (2025)
by: Golbert, Simon
Published: (2025)
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
by: Gao, Chen-Xiao, et al.
Published: (2023)
by: Gao, Chen-Xiao, et al.
Published: (2023)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
RAT: Retrieval-Augmented Transformer for Click-Through Rate Prediction
by: Li, Yushen, et al.
Published: (2024)
by: Li, Yushen, et al.
Published: (2024)
Learning without Global Backpropagation via Synergistic Information Distillation
by: Ye, Chenhao, et al.
Published: (2025)
by: Ye, Chenhao, et al.
Published: (2025)
Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
by: Zhang, Beining, et al.
Published: (2025)
by: Zhang, Beining, et al.
Published: (2025)
Do Transformer World Models Give Better Policy Gradients?
by: Ma, Michel, et al.
Published: (2024)
by: Ma, Michel, et al.
Published: (2024)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
by: Mao, Hangyu, et al.
Published: (2025)
by: Mao, Hangyu, et al.
Published: (2025)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
by: Xu, Zihang, et al.
Published: (2026)
by: Xu, Zihang, et al.
Published: (2026)
ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards
by: Li, Fanxing, et al.
Published: (2025)
by: Li, Fanxing, et al.
Published: (2025)
Vertical Symbolic Regression via Deep Policy Gradient
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Backpropagation-Free Metropolis-Adjusted Langevin Algorithm
by: Cobb, Adam D., et al.
Published: (2025)
by: Cobb, Adam D., et al.
Published: (2025)
Towards Green AI in Fine-tuning Large Language Models via Adaptive Backpropagation
by: Huang, Kai, et al.
Published: (2023)
by: Huang, Kai, et al.
Published: (2023)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
HKAN: Hierarchical Kolmogorov-Arnold Network without Backpropagation
by: Dudek, Grzegorz, et al.
Published: (2025)
by: Dudek, Grzegorz, et al.
Published: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Dynamic Spectral Backpropagation for Efficient Neural Network Training
by: Muthuraman, Mannmohan
Published: (2025)
by: Muthuraman, Mannmohan
Published: (2025)
Exploring the Performance of Perforated Backpropagation through Further Experiments
by: Brenner, Rorry, et al.
Published: (2025)
by: Brenner, Rorry, et al.
Published: (2025)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
by: Bai, Fengshuo, et al.
Published: (2024)
by: Bai, Fengshuo, et al.
Published: (2024)
Direct Soft-Policy Sampling via Langevin Dynamics
by: Ki, Donghyeon, et al.
Published: (2026)
by: Ki, Donghyeon, et al.
Published: (2026)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
by: Wei, Xing, et al.
Published: (2025)
by: Wei, Xing, et al.
Published: (2025)
Policy Gradient with Kernel Quadrature
by: Hayakawa, Satoshi, et al.
Published: (2023)
by: Hayakawa, Satoshi, et al.
Published: (2023)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
by: Krčo, Nataša, et al.
Published: (2026)
by: Krčo, Nataša, et al.
Published: (2026)
Similar Items
-
Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
by: Abdi, Hossein, et al.
Published: (2025) -
Gradient Regularized Natural Gradients
by: Dash, Satya Prakash, et al.
Published: (2026) -
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
by: Huo, Yingxiao, et al.
Published: (2026) -
GRADE: Replacing Policy Gradients with Backpropagation for LLM Alignment
by: Nel, Lukas Abrie
Published: (2025) -
Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
by: Flügel, Katharina, et al.
Published: (2024)