Action-Dependent Optimality-Preserving Reward Shaping
Fuente:
arXiv
Saved in:
| Main Authors: | Forbes, Grant C., Wang, Jianxun, Villalobos-Arias, Leonardo, Jhala, Arnav, Roberts, David L. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
Cost and Reward Infused Metric Elicitation
by: Bhateja, Chethan, et al.
Published: (2025)
by: Bhateja, Chethan, et al.
Published: (2025)
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
by: Quadros, André, et al.
Published: (2025)
by: Quadros, André, et al.
Published: (2025)
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
by: Lachapelle, Sébastien, et al.
Published: (2024)
by: Lachapelle, Sébastien, et al.
Published: (2024)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Exploring Neural Granger Causality with xLSTMs: Unveiling Temporal Dependencies in Complex Data
by: Poonia, Harsh, et al.
Published: (2025)
by: Poonia, Harsh, et al.
Published: (2025)
Graph Neural Network Based Action Ranking for Planning
by: Mangannavar, Rajesh, et al.
Published: (2024)
by: Mangannavar, Rajesh, et al.
Published: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Difference Rewards Policy Gradients
by: Castellini, Jacopo, et al.
Published: (2020)
by: Castellini, Jacopo, et al.
Published: (2020)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
by: Mongaras, Gabriel, et al.
Published: (2026)
by: Mongaras, Gabriel, et al.
Published: (2026)
Contextual Combinatorial Bandits with Changing Action Sets via Gaussian Processes
by: Nika, Andi, et al.
Published: (2021)
by: Nika, Andi, et al.
Published: (2021)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
by: Malarkkan, Arun Vignesh, et al.
Published: (2025)
When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning
by: Kujur, Arahan
Published: (2026)
by: Kujur, Arahan
Published: (2026)
Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPs
by: Hung, Wei, et al.
Published: (2025)
by: Hung, Wei, et al.
Published: (2025)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
Algebraic Machine Learning for Small-to-Medium Datasets Is Competitive against Strong Standard Baselines
by: Mendez, David, et al.
Published: (2026)
by: Mendez, David, et al.
Published: (2026)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
A Constraint-Preserving Neural Network Approach for Solving Mean-Field Games Equilibrium
by: Liu, Jinwei, et al.
Published: (2025)
by: Liu, Jinwei, et al.
Published: (2025)
How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures
by: Gupta, Krishnam
Published: (2026)
by: Gupta, Krishnam
Published: (2026)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
Data-Incremental Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2024)
by: Gai, Sibo, et al.
Published: (2024)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
by: T, Kevin Joshua, et al.
Published: (2024)
by: T, Kevin Joshua, et al.
Published: (2024)
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2023)
by: Gai, Sibo, et al.
Published: (2023)
Representation learning with CGAN for casual inference
by: Weng, Zhaotian, et al.
Published: (2024)
by: Weng, Zhaotian, et al.
Published: (2024)
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
by: Wang, Fang, et al.
Published: (2025)
by: Wang, Fang, et al.
Published: (2025)
Decoding Rewards in Competitive Games: Inverse Game Theory with Entropy Regularization
by: Liao, Junyi, et al.
Published: (2026)
by: Liao, Junyi, et al.
Published: (2026)
The Bayesian Confidence (BACON) Estimator for Deep Neural Networks
by: Kee, Patrick D., et al.
Published: (2024)
by: Kee, Patrick D., et al.
Published: (2024)
Explainable Graph Representation Learning via Graph Pattern Analysis
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
Improving the Expressiveness of $K$-hop Message-Passing GNNs by Injecting Contextualized Substructure Information
by: Yao, Tianjun, et al.
Published: (2024)
by: Yao, Tianjun, et al.
Published: (2024)
Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms
by: Avery, Katherine, et al.
Published: (2025)
by: Avery, Katherine, et al.
Published: (2025)
New Paradigm of Adversarial Training: Releasing Accuracy-Robustness Trade-Off via Dummy Class
by: Wang, Yanyun, et al.
Published: (2024)
by: Wang, Yanyun, et al.
Published: (2024)
Safe Reinforcement Learning with Preference-based Constraint Inference
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Graceful task adaptation with a bi-hemispheric RL agent
by: Nicholas, Grant, et al.
Published: (2024)
by: Nicholas, Grant, et al.
Published: (2024)
Similar Items
-
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024) -
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024) -
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025) -
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024) -
Cost and Reward Infused Metric Elicitation
by: Bhateja, Chethan, et al.
Published: (2025)