On the "Causality" Step in Policy Gradient Derivations: A Pedagogical Reconciliation of Full Return and Reward-to-Go
Fuente:
arXiv
Saved in:
| Main Author: | Siboni, Nima H. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Driven Heuristic Synthesis for Industrial Process Control: Lessons from Hot Steel Rolling
by: Siboni, Nima H., et al.
Published: (2026)
by: Siboni, Nima H., et al.
Published: (2026)
Decoupling Return-to-Go for Efficient Decision Transformer
by: Wang, Yongyi, et al.
Published: (2026)
by: Wang, Yongyi, et al.
Published: (2026)
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
by: Mead, Harry, et al.
Published: (2025)
by: Mead, Harry, et al.
Published: (2025)
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023)
by: Kraske, Benjamin D., et al.
Published: (2023)
Stabilizing Policy Gradient Methods via Reward Profiling
by: Ahmed, Shihab, et al.
Published: (2025)
by: Ahmed, Shihab, et al.
Published: (2025)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
by: Koo, Juil, et al.
Published: (2026)
by: Koo, Juil, et al.
Published: (2026)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue
by: Du, Huifang, et al.
Published: (2024)
by: Du, Huifang, et al.
Published: (2024)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Step-by-Step Causality: Transparent Causal Discovery with Multi-Agent Tree-Query and Adversarial Confidence Estimation
by: Ding, Ziyi, et al.
Published: (2026)
by: Ding, Ziyi, et al.
Published: (2026)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026)
by: Cho, Minjae, et al.
Published: (2026)
LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
by: Rezaei, Mohammad, et al.
Published: (2026)
by: Rezaei, Mohammad, et al.
Published: (2026)
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
by: Lyu, Guwen, et al.
Published: (2024)
by: Lyu, Guwen, et al.
Published: (2024)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
by: Jabs, Dennis, et al.
Published: (2026)
by: Jabs, Dennis, et al.
Published: (2026)
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Reward Models Identify Consistency, Not Causality
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
by: Reber, David, et al.
Published: (2024)
by: Reber, David, et al.
Published: (2024)
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
by: Li, Shaolong, et al.
Published: (2026)
by: Li, Shaolong, et al.
Published: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
by: Lee, Chi-Chang, et al.
Published: (2025)
by: Lee, Chi-Chang, et al.
Published: (2025)
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
by: Liu, Yuchen, et al.
Published: (2026)
by: Liu, Yuchen, et al.
Published: (2026)
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
by: Yoon, Deokgyu, et al.
Published: (2026)
by: Yoon, Deokgyu, et al.
Published: (2026)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
GOPO: Policy Optimization using Ranked Rewards
by: Choi, Kyuseong, et al.
Published: (2026)
by: Choi, Kyuseong, et al.
Published: (2026)
Gradient Step Plug-and-Play Model for Dental Cone-Beam CT Reconstruction
by: Tatachak, Idris, et al.
Published: (2026)
by: Tatachak, Idris, et al.
Published: (2026)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
GoRA: Gradient-driven Adaptive Low Rank Adaptation
by: He, Haonan, et al.
Published: (2025)
by: He, Haonan, et al.
Published: (2025)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
Learning Causally Invariant Reward Functions from Diverse Demonstrations
by: Ovinnikov, Ivan, et al.
Published: (2024)
by: Ovinnikov, Ivan, et al.
Published: (2024)
Similar Items
-
LLM-Driven Heuristic Synthesis for Industrial Process Control: Lessons from Hot Steel Rolling
by: Siboni, Nima H., et al.
Published: (2026) -
Decoupling Return-to-Go for Efficient Decision Transformer
by: Wang, Yongyi, et al.
Published: (2026) -
Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
by: Mead, Harry, et al.
Published: (2025) -
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023) -
Stabilizing Policy Gradient Methods via Reward Profiling
by: Ahmed, Shihab, et al.
Published: (2025)