On The Expressivity of Objective-Specification Formalisms in Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Subramani, Rohan, Williams, Marcus, Heitmann, Max, Holm, Halfdan, Griffin, Charlie, Skalse, Joar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022)
por: Skalse, Joar, et al.
Publicado: (2022)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
por: Fluri, Lukas, et al.
Publicado: (2024)
por: Fluri, Lukas, et al.
Publicado: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
por: Skalse, Joar, et al.
Publicado: (2023)
por: Skalse, Joar, et al.
Publicado: (2023)
Automating Formal Verification with Reinforcement Learning and Recursive Inference
por: Tan, Max
Publicado: (2026)
por: Tan, Max
Publicado: (2026)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
por: Williams, Kai, et al.
Publicado: (2025)
por: Williams, Kai, et al.
Publicado: (2025)
Multi-objective Reinforcement learning from AI Feedback
por: Williams, Marcus
Publicado: (2024)
por: Williams, Marcus
Publicado: (2024)
Scalable Multi-Objective Robot Reinforcement Learning through Gradient Conflict Resolution
por: Munn, Humphrey, et al.
Publicado: (2025)
por: Munn, Humphrey, et al.
Publicado: (2025)
EXPO: Stable Reinforcement Learning with Expressive Policies
por: Dong, Perry, et al.
Publicado: (2025)
por: Dong, Perry, et al.
Publicado: (2025)
Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
por: Surdej, Rafał, et al.
Publicado: (2025)
por: Surdej, Rafał, et al.
Publicado: (2025)
Towards Formalizing Reinforcement Learning Theory
por: Zhang, Shangtong
Publicado: (2025)
por: Zhang, Shangtong
Publicado: (2025)
Polychromic Objectives for Reinforcement Learning
por: Hamid, Jubayer Ibn, et al.
Publicado: (2025)
por: Hamid, Jubayer Ibn, et al.
Publicado: (2025)
Pareto Set Learning for Multi-Objective Reinforcement Learning
por: Liu, Erlong, et al.
Publicado: (2025)
por: Liu, Erlong, et al.
Publicado: (2025)
Preference-based Multi-Objective Reinforcement Learning
por: Mu, Ni, et al.
Publicado: (2025)
por: Mu, Ni, et al.
Publicado: (2025)
Objective-Specific Privileged Bases via Full-Prefix Matryoshka Learning
por: Talukder, Arghamitra, et al.
Publicado: (2026)
por: Talukder, Arghamitra, et al.
Publicado: (2026)
Towards Provable Emergence of In-Context Reinforcement Learning
por: Wang, Jiuqi, et al.
Publicado: (2025)
por: Wang, Jiuqi, et al.
Publicado: (2025)
Optimistic Reinforcement Learning with Quantile Objectives
por: Alipour-Vaezi, Mohammad, et al.
Publicado: (2025)
por: Alipour-Vaezi, Mohammad, et al.
Publicado: (2025)
On Generalization Across Environments In Multi-Objective Reinforcement Learning
por: Teoh, Jayden, et al.
Publicado: (2025)
por: Teoh, Jayden, et al.
Publicado: (2025)
Expressive Value Learning for Scalable Offline Reinforcement Learning
por: Espinosa-Dice, Nicolas, et al.
Publicado: (2025)
por: Espinosa-Dice, Nicolas, et al.
Publicado: (2025)
Expressive Temporal Specifications for Reward Monitoring
por: Adalat, Omar, et al.
Publicado: (2025)
por: Adalat, Omar, et al.
Publicado: (2025)
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
por: Doo, JaeHyeok, et al.
Publicado: (2026)
por: Doo, JaeHyeok, et al.
Publicado: (2026)
A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
por: Chen, Ying-Tu, et al.
Publicado: (2026)
por: Chen, Ying-Tu, et al.
Publicado: (2026)
Theoretical Study of Conflict-Avoidant Multi-Objective Reinforcement Learning
por: Wang, Yudan, et al.
Publicado: (2024)
por: Wang, Yudan, et al.
Publicado: (2024)
Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning
por: Park, Giseung, et al.
Publicado: (2025)
por: Park, Giseung, et al.
Publicado: (2025)
Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
por: Ganesh, Swetha, et al.
Publicado: (2026)
por: Ganesh, Swetha, et al.
Publicado: (2026)
Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion
por: Park, Giseung, et al.
Publicado: (2026)
por: Park, Giseung, et al.
Publicado: (2026)
Benchmarking Offline Multi-Objective Reinforcement Learning in Critical Care
por: Bansal, Aryaman, et al.
Publicado: (2025)
por: Bansal, Aryaman, et al.
Publicado: (2025)
Multi-Objective Reinforcement Learning for Generating Covalent Inhibitor Candidates
por: Gil, Renee
Publicado: (2026)
por: Gil, Renee
Publicado: (2026)
Reinforcement Learning with $ω$-Regular Objectives and Constraints
por: Wagner, Dominik, et al.
Publicado: (2025)
por: Wagner, Dominik, et al.
Publicado: (2025)
Multi-Objective Reinforcement Learning for Water Management
por: Osika, Zuzanna, et al.
Publicado: (2025)
por: Osika, Zuzanna, et al.
Publicado: (2025)
Demonstration Guided Multi-Objective Reinforcement Learning
por: Lu, Junlin, et al.
Publicado: (2024)
por: Lu, Junlin, et al.
Publicado: (2024)
The Formalism-Implementation Gap in Reinforcement Learning Research
por: Castro, Pablo Samuel
Publicado: (2025)
por: Castro, Pablo Samuel
Publicado: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
por: Griffin, Charlie, et al.
Publicado: (2024)
por: Griffin, Charlie, et al.
Publicado: (2024)
Effective Reward Specification in Deep Reinforcement Learning
por: Roy, Julien
Publicado: (2024)
por: Roy, Julien
Publicado: (2024)
Active Learning of Molecular Data for Task-Specific Objectives
por: Ghosh, Kunal, et al.
Publicado: (2024)
por: Ghosh, Kunal, et al.
Publicado: (2024)
Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data
por: Deng, Shilong, et al.
Publicado: (2025)
por: Deng, Shilong, et al.
Publicado: (2025)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
por: Xie, Zixuan, et al.
Publicado: (2026)
por: Xie, Zixuan, et al.
Publicado: (2026)
Ejemplares similares
-
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
por: Skalse, Joar, et al.
Publicado: (2024) -
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
por: Skalse, Joar, et al.
Publicado: (2024) -
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
por: Skalse, Joar, et al.
Publicado: (2024) -
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
por: Skalse, Joar, et al.
Publicado: (2024) -
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022)