Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
Fuente:
arXiv
Saved in:
| Main Authors: | Thrampoulidis, Christos, Mahdavi, Sadegh, Deng, Wenlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)
by: Mahdavi, Sadegh, et al.
Published: (2023)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
by: Lidayan, Aly, et al.
Published: (2024)
by: Lidayan, Aly, et al.
Published: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
by: Sun, Mingfei
Published: (2026)
by: Sun, Mingfei
Published: (2026)
From Graph Diffusion to Graph Classification
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
by: Xian, Jia Jun Cheng, et al.
Published: (2024)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)
by: Deng, Wenlong, et al.
Published: (2023)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
by: Xu, Zihang, et al.
Published: (2026)
by: Xu, Zihang, et al.
Published: (2026)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Stabilizing Policy Gradient Methods via Reward Profiling
by: Ahmed, Shihab, et al.
Published: (2025)
by: Ahmed, Shihab, et al.
Published: (2025)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
by: Holmes, Ian, et al.
Published: (2025)
by: Holmes, Ian, et al.
Published: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Maximize Your Diffusion: A Study into Reward Maximization and Alignment for Diffusion-based Control
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
by: Walder, Christian, et al.
Published: (2025)
by: Walder, Christian, et al.
Published: (2025)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
SURGE: Surrogate Gradient Adaptation in Binary Neural Networks
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
Maximally Permissive Reward Machines
by: Varricchione, Giovanni, et al.
Published: (2024)
by: Varricchione, Giovanni, et al.
Published: (2024)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026)
by: Cho, Minjae, et al.
Published: (2026)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Bayesian Deep Learning Via Expectation Maximization and Turbo Deep Approximate Message Passing
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
by: Zhang, Qixin, et al.
Published: (2024)
by: Zhang, Qixin, et al.
Published: (2024)
Harnessing Optimization Dynamics for Curvature-Informed Model Merging
by: Mahdavinia, Pouria, et al.
Published: (2025)
by: Mahdavinia, Pouria, et al.
Published: (2025)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Combining Automated Optimisation of Hyperparameters and Reward Shape
by: Dierkes, Julian, et al.
Published: (2024)
by: Dierkes, Julian, et al.
Published: (2024)
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
by: Gan, Wangjie, et al.
Published: (2026)
by: Gan, Wangjie, et al.
Published: (2026)
Learning Surrogates for Offline Black-Box Optimization via Gradient Matching
by: Hoang, Minh, et al.
Published: (2025)
by: Hoang, Minh, et al.
Published: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
by: Pavlenko, Kirill, et al.
Published: (2026)
by: Pavlenko, Kirill, et al.
Published: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Learning to Recover: Dynamic Reward Shaping with Wheel-Leg Coordination for Fallen Robots
by: Deng, Boyuan, et al.
Published: (2025)
by: Deng, Boyuan, et al.
Published: (2025)
Similar Items
-
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025) -
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023) -
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
by: Deng, Wenlong, et al.
Published: (2024) -
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
by: Deng, Wenlong, et al.
Published: (2024) -
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
by: Lidayan, Aly, et al.
Published: (2024)