Potential-Based Reward Shaping For Intrinsic Motivation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Forbes, Grant C., Gupta, Nitish, Villalobos-Arias, Leonardo, Potts, Colin M., Jhala, Arnav, Roberts, David L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
von: Forbes, Grant C., et al.
Veröffentlicht: (2024)
Action-Dependent Optimality-Preserving Reward Shaping
von: Forbes, Grant C., et al.
Veröffentlicht: (2025)
von: Forbes, Grant C., et al.
Veröffentlicht: (2025)
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
von: Quadros, André, et al.
Veröffentlicht: (2025)
von: Quadros, André, et al.
Veröffentlicht: (2025)
Fusing Rewards and Preferences in Reinforcement Learning
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors
von: Villalobos-Arias, Leonardo, et al.
Veröffentlicht: (2025)
von: Villalobos-Arias, Leonardo, et al.
Veröffentlicht: (2025)
Cost and Reward Infused Metric Elicitation
von: Bhateja, Chethan, et al.
Veröffentlicht: (2025)
von: Bhateja, Chethan, et al.
Veröffentlicht: (2025)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
von: Li, Bangzheng, et al.
Veröffentlicht: (2024)
von: Li, Bangzheng, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
von: Singha, Disha
Veröffentlicht: (2026)
von: Singha, Disha
Veröffentlicht: (2026)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
AGOP-IxG: A Gradient Covariance Filter for Local Feature Attribution on Tabular Data, with a Controlled Benchmark
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
Understanding Variational Autoencoders with Intrinsic Dimension and Information Imbalance
von: Camboulin, Charles, et al.
Veröffentlicht: (2024)
von: Camboulin, Charles, et al.
Veröffentlicht: (2024)
X-Factor: Quality Is a Dataset-Intrinsic Property
von: Couch, Josiah, et al.
Veröffentlicht: (2025)
von: Couch, Josiah, et al.
Veröffentlicht: (2025)
CoxSE: Exploring the Potential of Self-Explaining Neural Networks with Cox Proportional Hazards Model for Survival Analysis
von: Alabdallah, Abdallah, et al.
Veröffentlicht: (2024)
von: Alabdallah, Abdallah, et al.
Veröffentlicht: (2024)
Difference Rewards Policy Gradients
von: Castellini, Jacopo, et al.
Veröffentlicht: (2020)
von: Castellini, Jacopo, et al.
Veröffentlicht: (2020)
How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures
von: Gupta, Krishnam
Veröffentlicht: (2026)
von: Gupta, Krishnam
Veröffentlicht: (2026)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
von: Furuyama, Ryoma, et al.
Veröffentlicht: (2024)
von: Furuyama, Ryoma, et al.
Veröffentlicht: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
von: Yousaf, Iqra
Veröffentlicht: (2024)
von: Yousaf, Iqra
Veröffentlicht: (2024)
Prompting Neural-Guided Equation Discovery Based on Residuals
von: Brugger, Jannis, et al.
Veröffentlicht: (2025)
von: Brugger, Jannis, et al.
Veröffentlicht: (2025)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management
von: Zaregarizi, Shadmehr, et al.
Veröffentlicht: (2026)
von: Zaregarizi, Shadmehr, et al.
Veröffentlicht: (2026)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
von: Belcamino, Valerio, et al.
Veröffentlicht: (2026)
Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
von: Cardoso, Lucas, et al.
Veröffentlicht: (2025)
ES-C51: Expected Sarsa Based C51 Distributional Reinforcement Learning Algorithm
von: Tandon, Rijul, et al.
Veröffentlicht: (2025)
von: Tandon, Rijul, et al.
Veröffentlicht: (2025)
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2026)
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2026)
FluidWorld: Reaction-Diffusion Dynamics as a Predictive Substrate for World Models
von: Polly, Fabien
Veröffentlicht: (2026)
von: Polly, Fabien
Veröffentlicht: (2026)
I-GLIDE: Input Groups for Latent Health Indicators in Degradation Estimation
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
von: Thil, Lucas, et al.
Veröffentlicht: (2025)
Algebraic Machine Learning for Small-to-Medium Datasets Is Competitive against Strong Standard Baselines
von: Mendez, David, et al.
Veröffentlicht: (2026)
von: Mendez, David, et al.
Veröffentlicht: (2026)
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
von: Ribeiro, José, et al.
Veröffentlicht: (2022)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
von: Alsheikh, Ahmad, et al.
Veröffentlicht: (2025)
von: Alsheikh, Ahmad, et al.
Veröffentlicht: (2025)
The Bayesian Confidence (BACON) Estimator for Deep Neural Networks
von: Kee, Patrick D., et al.
Veröffentlicht: (2024)
von: Kee, Patrick D., et al.
Veröffentlicht: (2024)
Safe Reinforcement Learning with Preference-based Constraint Inference
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Graceful task adaptation with a bi-hemispheric RL agent
von: Nicholas, Grant, et al.
Veröffentlicht: (2024)
von: Nicholas, Grant, et al.
Veröffentlicht: (2024)
TrajGPT: Controlled Synthetic Trajectory Generation Using a Multitask Transformer-Based Spatiotemporal Model
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2024)
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2024)
Graph Neural Network Based Action Ranking for Planning
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
von: Forbes, Grant C., et al.
Veröffentlicht: (2024) -
Action-Dependent Optimality-Preserving Reward Shaping
von: Forbes, Grant C., et al.
Veröffentlicht: (2025) -
LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
von: Quadros, André, et al.
Veröffentlicht: (2025) -
Fusing Rewards and Preferences in Reinforcement Learning
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025) -
Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors
von: Villalobos-Arias, Leonardo, et al.
Veröffentlicht: (2025)