Bootstrapped Reward Shaping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adamczyk, Jacob, Makarenko, Volodymyr, Tiomkin, Stas, Kulkarni, Rahul V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024)
Maximum Entropy Exploration Without the Rollouts
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
Thermodynamics of Reinforcement Learning Curricula
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)
Multi-Resolution Diffusion for Privacy-Sensitive Recommender Systems
von: Lilienthal, Derek, et al.
Veröffentlicht: (2023)
von: Lilienthal, Derek, et al.
Veröffentlicht: (2023)
SuPLE: Robot Learning with Lyapunov Rewards
von: Nguyen, Phu, et al.
Veröffentlicht: (2024)
von: Nguyen, Phu, et al.
Veröffentlicht: (2024)
Exploration Behavior of Untrained Policies
von: Adamczyk, Jacob
Veröffentlicht: (2025)
von: Adamczyk, Jacob
Veröffentlicht: (2025)
Inferring Transition Dynamics from Value Functions
von: Adamczyk, Jacob
Veröffentlicht: (2025)
von: Adamczyk, Jacob
Veröffentlicht: (2025)
Emergence of Physical Intelligence via Controllable Information Production
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
Learning telic-controllable state representations
von: Amir, Nadav, et al.
Veröffentlicht: (2024)
von: Amir, Nadav, et al.
Veröffentlicht: (2024)
Goals and the Structure of Experience
von: Amir, Nadav, et al.
Veröffentlicht: (2025)
von: Amir, Nadav, et al.
Veröffentlicht: (2025)
Decentralized Traffic Flow Optimization Through Intrinsic Motivation
von: Papala, Himaja, et al.
Veröffentlicht: (2025)
von: Papala, Himaja, et al.
Veröffentlicht: (2025)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
von: Kar, Avik, et al.
Veröffentlicht: (2024)
von: Kar, Avik, et al.
Veröffentlicht: (2024)
Combining Automated Optimisation of Hyperparameters and Reward Shape
von: Dierkes, Julian, et al.
Veröffentlicht: (2024)
von: Dierkes, Julian, et al.
Veröffentlicht: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Automatic Reward Shaping from Confounded Offline Data
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge
von: Wu, Xiefeng
Veröffentlicht: (2024)
von: Wu, Xiefeng
Veröffentlicht: (2024)
Automatic Reward Shaping from Multi-Objective Human Heuristics
von: Xie, Yuqing, et al.
Veröffentlicht: (2025)
von: Xie, Yuqing, et al.
Veröffentlicht: (2025)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
von: Lin, Wenze, et al.
Veröffentlicht: (2026)
von: Lin, Wenze, et al.
Veröffentlicht: (2026)
Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
von: Okudo, Takato, et al.
Veröffentlicht: (2021)
von: Okudo, Takato, et al.
Veröffentlicht: (2021)
Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
von: Ma, Haozhe, et al.
Veröffentlicht: (2024)
von: Ma, Haozhe, et al.
Veröffentlicht: (2024)
Evaluating machine learning models for predicting pesticide toxicity to honey bees
von: Adamczyk, Jakub, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jakub, et al.
Veröffentlicht: (2025)
Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
von: Praski, Mateusz, et al.
Veröffentlicht: (2025)
von: Praski, Mateusz, et al.
Veröffentlicht: (2025)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
von: Wu, Boyuan
Veröffentlicht: (2025)
von: Wu, Boyuan
Veröffentlicht: (2025)
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
On the Sample Efficiency of Abstractions and Potential-Based Reward Shaping in Reinforcement Learning
von: Canonaco, Giuseppe, et al.
Veröffentlicht: (2024)
von: Canonaco, Giuseppe, et al.
Veröffentlicht: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
von: Wang, Haichuan, et al.
Veröffentlicht: (2026)
von: Wang, Haichuan, et al.
Veröffentlicht: (2026)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2024)
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2024)
Confounding Robust Continuous Control via Automatic Reward Shaping
von: Juliani, Mateo, et al.
Veröffentlicht: (2026)
von: Juliani, Mateo, et al.
Veröffentlicht: (2026)
Bootstrapping Expectiles in Reinforcement Learning
von: Clavier, Pierre, et al.
Veröffentlicht: (2024)
von: Clavier, Pierre, et al.
Veröffentlicht: (2024)
Imitation Bootstrapped Reinforcement Learning
von: Hu, Hengyuan, et al.
Veröffentlicht: (2023)
von: Hu, Hengyuan, et al.
Veröffentlicht: (2023)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
von: Xie, Tianbao, et al.
Veröffentlicht: (2023)
von: Xie, Tianbao, et al.
Veröffentlicht: (2023)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025) -
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025) -
Boosting Soft Q-Learning by Bounding
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2024) -
Maximum Entropy Exploration Without the Rollouts
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026) -
Thermodynamics of Reinforcement Learning Curricula
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2026)