BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
Fuente:
arXiv
Guardado en:
| Autores principales: | Lidayan, Aly, Dennis, Michael, Russell, Stuart |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Intrinsically-Motivated Humans and Agents in Open-World Exploration
por: Lidayan, Aly, et al.
Publicado: (2025)
por: Lidayan, Aly, et al.
Publicado: (2025)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
por: Lidayan, Aly, et al.
Publicado: (2025)
por: Lidayan, Aly, et al.
Publicado: (2025)
Bootstrapped Reward Shaping
por: Adamczyk, Jacob, et al.
Publicado: (2025)
por: Adamczyk, Jacob, et al.
Publicado: (2025)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
por: Thrampoulidis, Christos, et al.
Publicado: (2025)
por: Thrampoulidis, Christos, et al.
Publicado: (2025)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
por: Holmes, Ian, et al.
Publicado: (2025)
por: Holmes, Ian, et al.
Publicado: (2025)
From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge
por: Wu, Xiefeng
Publicado: (2024)
por: Wu, Xiefeng
Publicado: (2024)
Combining Automated Optimisation of Hyperparameters and Reward Shape
por: Dierkes, Julian, et al.
Publicado: (2024)
por: Dierkes, Julian, et al.
Publicado: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Automatic Reward Shaping from Confounded Offline Data
por: Li, Mingxuan, et al.
Publicado: (2025)
por: Li, Mingxuan, et al.
Publicado: (2025)
Potential-Based Reward Shaping For Intrinsic Motivation
por: Forbes, Grant C., et al.
Publicado: (2024)
por: Forbes, Grant C., et al.
Publicado: (2024)
Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
por: Ma, Haozhe, et al.
Publicado: (2024)
por: Ma, Haozhe, et al.
Publicado: (2024)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
por: Lin, Wenze, et al.
Publicado: (2026)
por: Lin, Wenze, et al.
Publicado: (2026)
Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
por: Okudo, Takato, et al.
Publicado: (2021)
por: Okudo, Takato, et al.
Publicado: (2021)
Automatic Reward Shaping from Multi-Objective Human Heuristics
por: Xie, Yuqing, et al.
Publicado: (2025)
por: Xie, Yuqing, et al.
Publicado: (2025)
NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
por: Liu, Lawrence, et al.
Publicado: (2025)
por: Liu, Lawrence, et al.
Publicado: (2025)
AI Alignment with Changing and Influenceable Reward Functions
por: Carroll, Micah, et al.
Publicado: (2024)
por: Carroll, Micah, et al.
Publicado: (2024)
On the Sample Efficiency of Abstractions and Potential-Based Reward Shaping in Reinforcement Learning
por: Canonaco, Giuseppe, et al.
Publicado: (2024)
por: Canonaco, Giuseppe, et al.
Publicado: (2024)
Reward Shaping for Inference-Time Alignment: A Stackelberg Game Perspective
por: Wang, Haichuan, et al.
Publicado: (2026)
por: Wang, Haichuan, et al.
Publicado: (2026)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
por: Wu, Boyuan
Publicado: (2025)
por: Wu, Boyuan
Publicado: (2025)
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
por: Li, Wenyun, et al.
Publicado: (2025)
por: Li, Wenyun, et al.
Publicado: (2025)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
por: Cho, Minjae, et al.
Publicado: (2026)
por: Cho, Minjae, et al.
Publicado: (2026)
Confounding Robust Continuous Control via Automatic Reward Shaping
por: Juliani, Mateo, et al.
Publicado: (2026)
por: Juliani, Mateo, et al.
Publicado: (2026)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
por: Bhambri, Siddhant, et al.
Publicado: (2024)
por: Bhambri, Siddhant, et al.
Publicado: (2024)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
por: Zhan, Simon Sinong, et al.
Publicado: (2024)
por: Zhan, Simon Sinong, et al.
Publicado: (2024)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
por: Liu, Yang, et al.
Publicado: (2026)
por: Liu, Yang, et al.
Publicado: (2026)
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
por: Xie, Tianbao, et al.
Publicado: (2023)
por: Xie, Tianbao, et al.
Publicado: (2023)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
por: Hugessen, Adriana, et al.
Publicado: (2024)
por: Hugessen, Adriana, et al.
Publicado: (2024)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
Learning to Recover: Dynamic Reward Shaping with Wheel-Leg Coordination for Fallen Robots
por: Deng, Boyuan, et al.
Publicado: (2025)
por: Deng, Boyuan, et al.
Publicado: (2025)
Fostering Intrinsic Motivation in Reinforcement Learning with Pretrained Foundation Models
por: Andres, Alain, et al.
Publicado: (2024)
por: Andres, Alain, et al.
Publicado: (2024)
A Generalized Acquisition Function for Preference-based Reward Learning
por: Ellis, Evan, et al.
Publicado: (2024)
por: Ellis, Evan, et al.
Publicado: (2024)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
por: Nguyen, Viet Bac, et al.
Publicado: (2026)
por: Nguyen, Viet Bac, et al.
Publicado: (2026)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
por: Yang, Daniel, et al.
Publicado: (2026)
por: Yang, Daniel, et al.
Publicado: (2026)
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs
por: Xie, Yutao, et al.
Publicado: (2026)
por: Xie, Yutao, et al.
Publicado: (2026)
Achieving Optimal Tissue Repair Through MARL with Reward Shaping and Curriculum Learning
por: Khan, Muhammad Al-Zafar, et al.
Publicado: (2025)
por: Khan, Muhammad Al-Zafar, et al.
Publicado: (2025)
Autotelic Agents with Intrinsically Motivated Goal-Conditioned Reinforcement Learning: a Short Survey
por: Colas, Cédric, et al.
Publicado: (2020)
por: Colas, Cédric, et al.
Publicado: (2020)
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
por: Liu, Zifan, et al.
Publicado: (2024)
por: Liu, Zifan, et al.
Publicado: (2024)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
por: Wang, Yibo, et al.
Publicado: (2024)
por: Wang, Yibo, et al.
Publicado: (2024)
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers
por: Guiducci, Leonardo, et al.
Publicado: (2025)
por: Guiducci, Leonardo, et al.
Publicado: (2025)
Ejemplares similares
-
Intrinsically-Motivated Humans and Agents in Open-World Exploration
por: Lidayan, Aly, et al.
Publicado: (2025) -
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
por: Lidayan, Aly, et al.
Publicado: (2025) -
Bootstrapped Reward Shaping
por: Adamczyk, Jacob, et al.
Publicado: (2025) -
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
por: Thrampoulidis, Christos, et al.
Publicado: (2025) -
Attention-Based Reward Shaping for Sparse and Delayed Rewards
por: Holmes, Ian, et al.
Publicado: (2025)