Sandbagging in a Simple Survival Bandit Problem
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dyer, Joel, Ornia, Daniel Jarne, Bishop, Nicholas, Calinescu, Anisoara, Wooldridge, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bayesian Decision Making around Experts
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025)
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025)
Automatic Differentiation of Agent-Based Models
von: Quera-Bofarull, Arnau, et al.
Veröffentlicht: (2025)
von: Quera-Bofarull, Arnau, et al.
Veröffentlicht: (2025)
Emergent Risk Awareness in Rational Agents under Resource Constraints
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025)
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025)
Causally Abstracted Multi-armed Bandits
von: Zennaro, Fabio Massimo, et al.
Veröffentlicht: (2024)
von: Zennaro, Fabio Massimo, et al.
Veröffentlicht: (2024)
Using causal abstractions to accelerate decision-making in complex bandit problems
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
A multi-objective combinatorial optimisation framework for large scale hierarchical population synthesis
von: Mahmood, Imran, et al.
Veröffentlicht: (2024)
von: Mahmood, Imran, et al.
Veröffentlicht: (2024)
Neural Network-Based Parameter Estimation of a Labour Market Agent-Based Model
von: Alves, M Lopes, et al.
Veröffentlicht: (2026)
von: Alves, M Lopes, et al.
Veröffentlicht: (2026)
A KL-regularization Framework for Learning to Plan with Adaptive Priors
von: Serra-Gomez, Álvaro, et al.
Veröffentlicht: (2025)
von: Serra-Gomez, Álvaro, et al.
Veröffentlicht: (2025)
Predictable Reinforcement Learning Dynamics through Entropy Rate Minimization
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2023)
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2023)
Removing Sandbagging in LLMs by Training with Weak Supervision
von: Ryd, Emil, et al.
Veröffentlicht: (2026)
von: Ryd, Emil, et al.
Veröffentlicht: (2026)
Robust Uncertainty Quantification Using Conformalised Monte Carlo Prediction
von: Bethell, Daniel, et al.
Veröffentlicht: (2023)
von: Bethell, Daniel, et al.
Veröffentlicht: (2023)
SynEHRgy: Synthesizing Mixed-Type Structured Electronic Health Records using Decoder-Only Transformers
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding
von: Bethell, Daniel, et al.
Veröffentlicht: (2024)
von: Bethell, Daniel, et al.
Veröffentlicht: (2024)
FeatEHR-LLM: Leveraging Large Language Models for Feature Engineering in Electronic Health Records
von: Karami, Hojjat, et al.
Veröffentlicht: (2026)
von: Karami, Hojjat, et al.
Veröffentlicht: (2026)
Fixed Point Explainability
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
Learning to Manage Investment Portfolios beyond Simple Utility Functions
von: Scholl, Maarten P., et al.
Veröffentlicht: (2025)
von: Scholl, Maarten P., et al.
Veröffentlicht: (2025)
Conformal Prediction for Dose-Response Models with Continuous Treatments
von: Verhaeghe, Jarne, et al.
Veröffentlicht: (2024)
von: Verhaeghe, Jarne, et al.
Veröffentlicht: (2024)
Partition Tree Weighting for Non-Stationary Stochastic Bandits
von: Veness, Joel, et al.
Veröffentlicht: (2025)
von: Veness, Joel, et al.
Veröffentlicht: (2025)
End-to-end PDDL Planning with Hardcoded and Dynamic Agents
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
Fair Algorithms with Probing for Multi-Agent Multi-Armed Bandits
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
Auditing Games for Sandbagging
von: Taylor, Jordan, et al.
Veröffentlicht: (2025)
von: Taylor, Jordan, et al.
Veröffentlicht: (2025)
Selective Reviews of Bandit Problems in AI via a Statistical View
von: Zhou, Pengjie, et al.
Veröffentlicht: (2024)
von: Zhou, Pengjie, et al.
Veröffentlicht: (2024)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
Incentivized Lipschitz Bandits
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2025)
Impatient Bandits: Optimizing for the Long-Term Without Delay
von: Zhang, Kelly W., et al.
Veröffentlicht: (2025)
von: Zhang, Kelly W., et al.
Veröffentlicht: (2025)
Survival Analysis with Adversarial Regularization
von: Potter, Michael, et al.
Veröffentlicht: (2023)
von: Potter, Michael, et al.
Veröffentlicht: (2023)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
Online Clustering of Dueling Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
Tree Ensembles for Contextual Bandits
von: Nilsson, Hannes, et al.
Veröffentlicht: (2024)
von: Nilsson, Hannes, et al.
Veröffentlicht: (2024)
Flickering Multi-Armed Bandits
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2026)
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2026)
Code Simulation as a Proxy for High-order Tasks in Large Language Models
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2025)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
von: Erez, Liad, et al.
Veröffentlicht: (2025)
von: Erez, Liad, et al.
Veröffentlicht: (2025)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
Tokenized Bandit for LLM Decoding and Alignment
von: Shin, Suho, et al.
Veröffentlicht: (2025)
von: Shin, Suho, et al.
Veröffentlicht: (2025)
Deceptive Exploration in Multi-armed Bandits
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
Causal Contextual Bandits with Adaptive Context
von: Madhavan, Rahul, et al.
Veröffentlicht: (2024)
von: Madhavan, Rahul, et al.
Veröffentlicht: (2024)
Neural Active Learning Beyond Bandits
von: Ban, Yikun, et al.
Veröffentlicht: (2024)
von: Ban, Yikun, et al.
Veröffentlicht: (2024)
Diffusion Models Meet Contextual Bandits
von: Aouali, Imad
Veröffentlicht: (2024)
von: Aouali, Imad
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bayesian Decision Making around Experts
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025) -
Automatic Differentiation of Agent-Based Models
von: Quera-Bofarull, Arnau, et al.
Veröffentlicht: (2025) -
Emergent Risk Awareness in Rational Agents under Resource Constraints
von: Ornia, Daniel Jarne, et al.
Veröffentlicht: (2025) -
Causally Abstracted Multi-armed Bandits
von: Zennaro, Fabio Massimo, et al.
Veröffentlicht: (2024) -
Using causal abstractions to accelerate decision-making in complex bandit problems
von: Dyer, Joel, et al.
Veröffentlicht: (2025)