Learning Rules from Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Puebla, Guillermo, Doumas, Leonidas A. A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Human-like generalization in a machine through predicate learning
por: Doumas, Leonidas A. A., et al.
Publicado: (2018)
por: Doumas, Leonidas A. A., et al.
Publicado: (2018)
A Theory of Relation Learning and Cross-domain Generalization
por: Doumas, Leonidas A. A., et al.
Publicado: (2019)
por: Doumas, Leonidas A. A., et al.
Publicado: (2019)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
por: Vegner, Ivan, et al.
Publicado: (2025)
por: Vegner, Ivan, et al.
Publicado: (2025)
Rule Based Rewards for Language Model Safety
por: Mu, Tong, et al.
Publicado: (2024)
por: Mu, Tong, et al.
Publicado: (2024)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
por: Wang, Tevin, et al.
Publicado: (2025)
por: Wang, Tevin, et al.
Publicado: (2025)
The Impact of Machine Learning Uncertainty on the Robustness of Counterfactual Explanations
por: Christodoulou, Leonidas, et al.
Publicado: (2026)
por: Christodoulou, Leonidas, et al.
Publicado: (2026)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
por: Peng, Hao, et al.
Publicado: (2026)
por: Peng, Hao, et al.
Publicado: (2026)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
por: Fanconi, Claudio, et al.
Publicado: (2025)
por: Fanconi, Claudio, et al.
Publicado: (2025)
FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models
por: Uluşan, Zeynel A., et al.
Publicado: (2026)
por: Uluşan, Zeynel A., et al.
Publicado: (2026)
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
por: Wang, Linji, et al.
Publicado: (2025)
por: Wang, Linji, et al.
Publicado: (2025)
Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
por: Infante, Guillermo, et al.
Publicado: (2024)
por: Infante, Guillermo, et al.
Publicado: (2024)
Reward Hacking in Rubric-Based Reinforcement Learning
por: Mahmoud, Anas, et al.
Publicado: (2026)
por: Mahmoud, Anas, et al.
Publicado: (2026)
RLNVR: Reinforcement Learning from Non-Verified Real-World Rewards
por: Krishnan, Rohit, et al.
Publicado: (2025)
por: Krishnan, Rohit, et al.
Publicado: (2025)
Learning Logical Rules using Minimum Message Length
por: Sharma, Ruben, et al.
Publicado: (2025)
por: Sharma, Ruben, et al.
Publicado: (2025)
Learned-Rule-Augmented Large Language Model Evaluators
por: Meng, Jie, et al.
Publicado: (2025)
por: Meng, Jie, et al.
Publicado: (2025)
Comparing Reinforcement Learning and Human Learning using the Game of Hidden Rules
por: Pulick, Eric, et al.
Publicado: (2023)
por: Pulick, Eric, et al.
Publicado: (2023)
Reward Learning from Multiple Feedback Types
por: Metz, Yannick, et al.
Publicado: (2025)
por: Metz, Yannick, et al.
Publicado: (2025)
RLSR: Reinforcement Learning from Self Reward
por: Simonds, Toby, et al.
Publicado: (2025)
por: Simonds, Toby, et al.
Publicado: (2025)
A Novel Framework for Uncertainty-Driven Adaptive Exploration
por: Bakopoulos, Leonidas, et al.
Publicado: (2025)
por: Bakopoulos, Leonidas, et al.
Publicado: (2025)
Semantic Association Rule Learning from Time Series Data and Knowledge Graphs
por: Karabulut, Erkan, et al.
Publicado: (2023)
por: Karabulut, Erkan, et al.
Publicado: (2023)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
por: Zhou, Zhiyuan, et al.
Publicado: (2022)
por: Zhou, Zhiyuan, et al.
Publicado: (2022)
SemiReward: A General Reward Model for Semi-supervised Learning
por: Li, Siyuan, et al.
Publicado: (2023)
por: Li, Siyuan, et al.
Publicado: (2023)
VIRAL: Vision-grounded Integration for Reward design And Learning
por: Cuzin-Rambaud, Valentin, et al.
Publicado: (2025)
por: Cuzin-Rambaud, Valentin, et al.
Publicado: (2025)
Rule by Rule: Learning with Confidence through Vocabulary Expansion
por: Nössig, Albert, et al.
Publicado: (2024)
por: Nössig, Albert, et al.
Publicado: (2024)
A Harmonic Mean Formulation of Average Reward Reinforcement Learning in SMDPs
por: Shtossel, Erel, et al.
Publicado: (2026)
por: Shtossel, Erel, et al.
Publicado: (2026)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
por: Kim, Kihyun, et al.
Publicado: (2024)
por: Kim, Kihyun, et al.
Publicado: (2024)
Pushdown Reward Machines for Reinforcement Learning
por: Varricchione, Giovanni, et al.
Publicado: (2025)
por: Varricchione, Giovanni, et al.
Publicado: (2025)
Discriminative Rule Learning for Outcome-Guided Process Model Discovery
por: Norouzifar, Ali, et al.
Publicado: (2025)
por: Norouzifar, Ali, et al.
Publicado: (2025)
Notes on the Reward Representation of Posterior Updates
por: Ortega, Pedro A.
Publicado: (2026)
por: Ortega, Pedro A.
Publicado: (2026)
Understanding Expressivity of GNN in Rule Learning
por: Qiu, Haiquan, et al.
Publicado: (2023)
por: Qiu, Haiquan, et al.
Publicado: (2023)
Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
por: Ishihara, Yu, et al.
Publicado: (2025)
por: Ishihara, Yu, et al.
Publicado: (2025)
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning?
por: Shihab, Ibne Farabi, et al.
Publicado: (2025)
por: Shihab, Ibne Farabi, et al.
Publicado: (2025)
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning
por: Abboud, Elie, et al.
Publicado: (2026)
por: Abboud, Elie, et al.
Publicado: (2026)
Confidence as a Reward: Transforming LLMs into Reward Models
por: Du, He, et al.
Publicado: (2025)
por: Du, He, et al.
Publicado: (2025)
Learning Robust Reward Machines from Noisy Labels
por: Parac, Roko, et al.
Publicado: (2024)
por: Parac, Roko, et al.
Publicado: (2024)
Learning Safety Constraints from Demonstrations with Unknown Rewards
por: Lindner, David, et al.
Publicado: (2023)
por: Lindner, David, et al.
Publicado: (2023)
Multi-Task Reward Learning from Human Ratings
por: Wu, Mingkang, et al.
Publicado: (2025)
por: Wu, Mingkang, et al.
Publicado: (2025)
SuPLE: Robot Learning with Lyapunov Rewards
por: Nguyen, Phu, et al.
Publicado: (2024)
por: Nguyen, Phu, et al.
Publicado: (2024)
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
por: Wang, Miao, et al.
Publicado: (2026)
por: Wang, Miao, et al.
Publicado: (2026)
Ejemplares similares
-
Human-like generalization in a machine through predicate learning
por: Doumas, Leonidas A. A., et al.
Publicado: (2018) -
A Theory of Relation Learning and Cross-domain Generalization
por: Doumas, Leonidas A. A., et al.
Publicado: (2019) -
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
por: Vegner, Ivan, et al.
Publicado: (2025) -
Rule Based Rewards for Language Model Safety
por: Mu, Tong, et al.
Publicado: (2024) -
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
por: Wang, Tevin, et al.
Publicado: (2025)