LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Quadros, André, Silva, Cassio, Alves, Ronnie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024)
by: Forbes, Grant C., et al.
Published: (2024)
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness
by: Cardoso, Lucas, et al.
Published: (2025)
by: Cardoso, Lucas, et al.
Published: (2025)
Beyond Random Sampling: Instance Quality-Based Data Partitioning via Item Response Theory
by: Cardoso, Lucas, et al.
Published: (2025)
by: Cardoso, Lucas, et al.
Published: (2025)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Standing on the shoulders of giants
by: Cardoso, Lucas Felipe Ferraro, et al.
Published: (2024)
by: Cardoso, Lucas Felipe Ferraro, et al.
Published: (2024)
Rewarded Region Replay (R3) for Policy Learning with Discrete Action Space
by: Li, Bangzheng, et al.
Published: (2024)
by: Li, Bangzheng, et al.
Published: (2024)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
Cost and Reward Infused Metric Elicitation
by: Bhateja, Chethan, et al.
Published: (2025)
by: Bhateja, Chethan, et al.
Published: (2025)
Action-Dependent Optimality-Preserving Reward Shaping
by: Forbes, Grant C., et al.
Published: (2025)
by: Forbes, Grant C., et al.
Published: (2025)
Hierarchical Reinforcement Learning with Targeted Causal Interventions
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
Data-Incremental Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2024)
by: Gai, Sibo, et al.
Published: (2024)
A Survey of Reinforcement Learning from Human Feedback
by: Kaufmann, Timo, et al.
Published: (2023)
by: Kaufmann, Timo, et al.
Published: (2023)
Perfecting Aircraft Maneuvers with Reinforcement Learning
by: Cilan, Atahan, et al.
Published: (2026)
by: Cilan, Atahan, et al.
Published: (2026)
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
by: Gai, Sibo, et al.
Published: (2023)
by: Gai, Sibo, et al.
Published: (2023)
A Comparative Analysis of Reinforcement Learning and Conventional Deep Learning Approaches for Bearing Fault Diagnosis
by: Çakır, Efe, et al.
Published: (2025)
by: Çakır, Efe, et al.
Published: (2025)
Playing Hex and Counter Wargames using Reinforcement Learning and Recurrent Neural Networks
by: Palma, Guilherme, et al.
Published: (2025)
by: Palma, Guilherme, et al.
Published: (2025)
Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
by: Fraser, Sandy, et al.
Published: (2025)
by: Fraser, Sandy, et al.
Published: (2025)
How Reliable and Stable are Explanations of XAI Methods?
by: Ribeiro, José, et al.
Published: (2024)
by: Ribeiro, José, et al.
Published: (2024)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
ES-C51: Expected Sarsa Based C51 Distributional Reinforcement Learning Algorithm
by: Tandon, Rijul, et al.
Published: (2025)
by: Tandon, Rijul, et al.
Published: (2025)
Fractional Policy Gradients: Reinforcement Learning with Long-Term Memory
by: Pawar, Urvi, et al.
Published: (2025)
by: Pawar, Urvi, et al.
Published: (2025)
Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement Learning
by: Li, Xinran, et al.
Published: (2024)
by: Li, Xinran, et al.
Published: (2024)
Distributional Reinforcement Learning for Condition-Based Maintenance of Multi-Pump Equipment
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
by: Chapman, James, et al.
Published: (2025)
by: Chapman, James, et al.
Published: (2025)
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
by: Tomashevskiy, Timofey
Published: (2026)
by: Tomashevskiy, Timofey
Published: (2026)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
by: Chen, Wen-Tse, et al.
Published: (2024)
by: Chen, Wen-Tse, et al.
Published: (2024)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
by: Yuan, Xin, et al.
Published: (2025)
by: Yuan, Xin, et al.
Published: (2025)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026)
by: Lan, Guangchen, et al.
Published: (2026)
FlowRL: Flow-Augmented Few-Shot Reinforcement Learning for Semi-Structured Sensor Data
by: Pivezhandi, Mohammad, et al.
Published: (2024)
by: Pivezhandi, Mohammad, et al.
Published: (2024)
SQARL: A Size-Agnostic Reinforcement Learning approach for Circuit Allocation in Distributed Quantum Architectures
by: Carballo, Víctor, et al.
Published: (2026)
by: Carballo, Víctor, et al.
Published: (2026)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
by: Lachapelle, Sébastien, et al.
Published: (2024)
by: Lachapelle, Sébastien, et al.
Published: (2024)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Similar Items
-
Potential-Based Reward Shaping For Intrinsic Motivation
by: Forbes, Grant C., et al.
Published: (2024) -
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025) -
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
by: Forbes, Grant C., et al.
Published: (2024) -
Explanations Based on Item Response Theory (eXirt): A Model-Specific Method to Explain Tree-Ensemble Model in Trust Perspective
by: Ribeiro, José, et al.
Published: (2022) -
Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness
by: Cardoso, Lucas, et al.
Published: (2025)