The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fluri, Lukas, Lang, Leon, Abate, Alessandro, Forré, Patrick, Krueger, David, Skalse, Joar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
von: Skalse, Joar, et al.
Veröffentlicht: (2023)
von: Skalse, Joar, et al.
Veröffentlicht: (2023)
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Modeling Human Beliefs about AI Behavior for Scalable Oversight
von: Lang, Leon, et al.
Veröffentlicht: (2025)
von: Lang, Leon, et al.
Veröffentlicht: (2025)
Defining and Characterizing Reward Hacking
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
von: Dalrymple, David "davidad", et al.
Veröffentlicht: (2024)
von: Dalrymple, David "davidad", et al.
Veröffentlicht: (2024)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
Provably Efficient Exploration in Reward Machines with Low Regret
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
Multivector Neurons: Better and Faster O(n)-Equivariant Clifford Graph Neural Networks
von: Liu, Cong, et al.
Veröffentlicht: (2024)
von: Liu, Cong, et al.
Veröffentlicht: (2024)
Debiasing Reward Models by Representation Learning with Guarantees
von: Ng, Ignavier, et al.
Veröffentlicht: (2025)
von: Ng, Ignavier, et al.
Veröffentlicht: (2025)
Networked Communication for Mean-Field Games with Function Approximation and Empirical Mean-Field Estimation
von: Benjamin, Patrick, et al.
Veröffentlicht: (2024)
von: Benjamin, Patrick, et al.
Veröffentlicht: (2024)
Clifford Group Equivariant Simplicial Message Passing Networks
von: Liu, Cong, et al.
Veröffentlicht: (2024)
von: Liu, Cong, et al.
Veröffentlicht: (2024)
Efficient Solution and Learning of Robust Factored MDPs
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
von: Vakili, Sattar, et al.
Veröffentlicht: (2024)
von: Vakili, Sattar, et al.
Veröffentlicht: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
Networked Communication for Decentralised Agents in Mean-Field Games
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
Clifford Group Equivariant Diffusion Models for 3D Molecular Generation
von: Liu, Cong, et al.
Veröffentlicht: (2025)
von: Liu, Cong, et al.
Veröffentlicht: (2025)
Data-Driven Online Model Selection With Regret Guarantees
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2023)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
FOSSIL: Regret-Minimizing Curriculum Learning for Metadata-Free and Low-Data Mpox Diagnosis
von: Han, Sahng-Min, et al.
Veröffentlicht: (2025)
von: Han, Sahng-Min, et al.
Veröffentlicht: (2025)
Strategic Arms with Side Communication Prevail Over Low-Regret MAB Algorithms
von: Yahmed, Ahmed Ben, et al.
Veröffentlicht: (2024)
von: Yahmed, Ahmed Ben, et al.
Veröffentlicht: (2024)
Neural Proofs for Sound Verification and Control of Complex Systems
von: Abate, Alessandro
Veröffentlicht: (2025)
von: Abate, Alessandro
Veröffentlicht: (2025)
Latent Representation and Simulation of Markov Processes via Time-Lagged Information Bottleneck
von: Federici, Marco, et al.
Veröffentlicht: (2023)
von: Federici, Marco, et al.
Veröffentlicht: (2023)
Temporal-Difference Variational Continual Learning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Clifford-Steerable Convolutional Neural Networks
von: Zhdanov, Maksim, et al.
Veröffentlicht: (2024)
von: Zhdanov, Maksim, et al.
Veröffentlicht: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
Early-Exit Neural Networks with Nested Prediction Sets
von: Jazbec, Metod, et al.
Veröffentlicht: (2023)
von: Jazbec, Metod, et al.
Veröffentlicht: (2023)
Certifiably Robust Policies for Uncertain Parametric Environments
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
AdS-GNN -- a Conformally Equivariant Graph Neural Network
von: Zhdanov, Maksim, et al.
Veröffentlicht: (2025)
von: Zhdanov, Maksim, et al.
Veröffentlicht: (2025)
Reward Bound for Behavioral Guarantee of Model-based Planning Agents
von: An, Zhiyu, et al.
Veröffentlicht: (2024)
von: An, Zhiyu, et al.
Veröffentlicht: (2024)
Goal Kernel Planning: Linearly-Solvable Non-Markovian Policies for Logical Tasks with Goal-Conditioned Options
von: Ringstrom, Thomas J., et al.
Veröffentlicht: (2020)
von: Ringstrom, Thomas J., et al.
Veröffentlicht: (2020)
Distributional Training Data Attribution: What do Influence Functions Sample?
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
von: Mlodozeniec, Bruno, et al.
Veröffentlicht: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
von: Ye, Hao, et al.
Veröffentlicht: (2026)
von: Ye, Hao, et al.
Veröffentlicht: (2026)
Reward Learning through Ranking Mean Squared Error
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2026)
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2026)
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
von: Wang, Linji, et al.
Veröffentlicht: (2025)
von: Wang, Linji, et al.
Veröffentlicht: (2025)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
von: Skalse, Joar, et al.
Veröffentlicht: (2024) -
STARC: A General Framework For Quantifying Differences Between Reward Functions
von: Skalse, Joar, et al.
Veröffentlicht: (2023) -
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
von: Skalse, Joar, et al.
Veröffentlicht: (2024)