Mitigating Goal Misgeneralization via Minimax Regret
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sadek, Karim Abdel, Farrugia-Roberts, Matthew, Anwar, Usman, Erlebach, Hannah, de Witt, Christian Schroeder, Krueger, David, Dennis, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward Model Ensembles Help Mitigate Overoptimization
von: Coste, Thomas, et al.
Veröffentlicht: (2023)
von: Coste, Thomas, et al.
Veröffentlicht: (2023)
Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization
von: Barj, Houda Nait El, et al.
Veröffentlicht: (2024)
von: Barj, Houda Nait El, et al.
Veröffentlicht: (2024)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
von: Elliott, Chris, et al.
Veröffentlicht: (2026)
von: Elliott, Chris, et al.
Veröffentlicht: (2026)
Structure and Scale in Simplicial Sequence Modelling
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2026)
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2026)
Getting By Goal Misgeneralization With a Little Help From a Mentor
von: Trinh, Tu, et al.
Veröffentlicht: (2024)
von: Trinh, Tu, et al.
Veröffentlicht: (2024)
Proximity to Losslessly Compressible Parameters
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2023)
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2023)
Refining Minimax Regret for Unsupervised Environment Design
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
von: Beukman, Michael, et al.
Veröffentlicht: (2024)
Algorithms for Caching and MTS with reduced number of predictions
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2024)
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2024)
Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks
von: Brumley, Madeline, et al.
Veröffentlicht: (2024)
von: Brumley, Madeline, et al.
Veröffentlicht: (2024)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
von: Bush, Thomas, et al.
Veröffentlicht: (2025)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
On the Minimax Regret in Online Ranking with Top-k Feedback
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2023)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
Minimax Regret Learning for Data with Heterogeneous Subgroups
von: Mo, Weibin, et al.
Veröffentlicht: (2024)
von: Mo, Weibin, et al.
Veröffentlicht: (2024)
Learning the Preferences of a Learning Agent
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2026)
von: Sadek, Karim Abdel, et al.
Veröffentlicht: (2026)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
von: Carroll, Liam, et al.
Veröffentlicht: (2025)
von: Carroll, Liam, et al.
Veröffentlicht: (2025)
Best of Both Worlds: Regret Minimization versus Minimax Play
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
Geometric Preference Elicitation for Minimax Regret Optimization in Uncertainty Matroids
von: Ellendula, Aditya Sai, et al.
Veröffentlicht: (2025)
von: Ellendula, Aditya Sai, et al.
Veröffentlicht: (2025)
Minimax-Optimal Policy Regret in Partially Observable Markov Games
von: Arora, Raman
Veröffentlicht: (2026)
von: Arora, Raman
Veröffentlicht: (2026)
Learning to Forget using Hypernetworks
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
von: Rangel, Jose Miguel Lara, et al.
Veröffentlicht: (2024)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
Minimax Regret Estimation for Generalizing Heterogeneous Treatment Effects with Multisite Data
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Mirror Learning: A Unifying Framework of Policy Optimisation
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2022)
von: Kuba, Jakub Grudzien, et al.
Veröffentlicht: (2022)
Dynamic Vocabulary Pruning in Early-Exit LLMs
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
Information-Theoretic Minimax Regret Bounds for Reinforcement Learning based on Duality
von: Bongole, Raghav, et al.
Veröffentlicht: (2024)
von: Bongole, Raghav, et al.
Veröffentlicht: (2024)
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
von: Di, Qiwei, et al.
Veröffentlicht: (2024)
von: Di, Qiwei, et al.
Veröffentlicht: (2024)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
von: Venhoff, Constantin, et al.
Veröffentlicht: (2024)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2024)
Bayesian Exploration Networks
von: Fellows, Mattie, et al.
Veröffentlicht: (2023)
von: Fellows, Mattie, et al.
Veröffentlicht: (2023)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
Temporal Task Diversity: Inductive Biases Under Non-Stationarity in Synthetic Sequence Modelling
von: Aswadi, Afiq Abdillah Effiezal, et al.
Veröffentlicht: (2026)
von: Aswadi, Afiq Abdillah Effiezal, et al.
Veröffentlicht: (2026)
Minimax Optimal Simple Regret in Two-Armed Best-Arm Identification
von: Kato, Masahiro
Veröffentlicht: (2024)
von: Kato, Masahiro
Veröffentlicht: (2024)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring
von: Anwar, Usman, et al.
Veröffentlicht: (2026)
von: Anwar, Usman, et al.
Veröffentlicht: (2026)
Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance
von: Zhao, Hangyi
Veröffentlicht: (2026)
von: Zhao, Hangyi
Veröffentlicht: (2026)
IDs for AI Systems
von: Chan, Alan, et al.
Veröffentlicht: (2024)
von: Chan, Alan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reward Model Ensembles Help Mitigate Overoptimization
von: Coste, Thomas, et al.
Veröffentlicht: (2023) -
Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization
von: Barj, Houda Nait El, et al.
Veröffentlicht: (2024) -
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
von: Elliott, Chris, et al.
Veröffentlicht: (2026) -
Structure and Scale in Simplicial Sequence Modelling
von: Farrugia-Roberts, Matthew
Veröffentlicht: (2026) -
Getting By Goal Misgeneralization With a Little Help From a Mentor
von: Trinh, Tu, et al.
Veröffentlicht: (2024)