Auditing Games for Sandbagging
Fuente:
arXiv
Saved in:
| Main Authors: | Taylor, Jordan, Black, Sid, Bowen, Dillon, Read, Thomas, Golechha, Satvik, Zelenka-Martin, Alex, Makins, Oliver, Kissane, Connor, Ayonrinde, Kola, Merizian, Jacob, Marks, Samuel, Cundy, Chris, Bloom, Joseph |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024)
by: Ayonrinde, Kola
Published: (2024)
Propensity Inference: Environmental Contributors to LLM Behaviour
by: Järviniemi, Olli, et al.
Published: (2026)
by: Järviniemi, Olli, et al.
Published: (2026)
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024)
by: Golechha, Satvik
Published: (2024)
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025)
by: Ayonrinde, Kola, et al.
Published: (2025)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
by: Ayonrinde, Kola, et al.
Published: (2024)
by: Ayonrinde, Kola, et al.
Published: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
by: Natarajan, Vikram, et al.
Published: (2026)
by: Natarajan, Vikram, et al.
Published: (2026)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
by: Balappanawar, Ishwar, et al.
Published: (2025)
by: Balappanawar, Ishwar, et al.
Published: (2025)
Training Neural Networks for Modularity aids Interpretability
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
by: Chanin, David, et al.
Published: (2024)
by: Chanin, David, et al.
Published: (2024)
NICE: To Optimize In-Context Examples or Not?
by: Srivastava, Pragya, et al.
Published: (2024)
by: Srivastava, Pragya, et al.
Published: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
by: Kissane, Connor, et al.
Published: (2024)
by: Kissane, Connor, et al.
Published: (2024)
From Mechanistic to Compositional Interpretability
by: Gauderis, Ward, et al.
Published: (2026)
by: Gauderis, Ward, et al.
Published: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
by: Lidayan, Aly, et al.
Published: (2025)
by: Lidayan, Aly, et al.
Published: (2025)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
by: Cundy, Chris, et al.
Published: (2025)
by: Cundy, Chris, et al.
Published: (2025)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
by: Cundy, Chris, et al.
Published: (2023)
by: Cundy, Chris, et al.
Published: (2023)
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration
by: Yadav, Advait, et al.
Published: (2026)
by: Yadav, Advait, et al.
Published: (2026)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
by: Tice, Cameron, et al.
Published: (2024)
by: Tice, Cameron, et al.
Published: (2024)
Do Large Language Models Know What They Are Capable Of?
by: Barkan, Casey O., et al.
Published: (2025)
by: Barkan, Casey O., et al.
Published: (2025)
Sandbagging in a Simple Survival Bandit Problem
by: Dyer, Joel, et al.
Published: (2025)
by: Dyer, Joel, et al.
Published: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
by: Ryd, Emil, et al.
Published: (2026)
by: Ryd, Emil, et al.
Published: (2026)
Algunas observaciones sobre los regímenes especiales de seguridad social en la agricultura
by: A. Zelenka
Published: (1963)
by: A. Zelenka
Published: (1963)
Some remarks on special social security schemes for agriculture
by: A. Zelenka
Published: (1963)
by: A. Zelenka
Published: (1963)
Quelques remarques sur les régimes spéciaux de sécurité sociale des agriculteurs
by: A. Zelenka
Published: (1963)
by: A. Zelenka
Published: (1963)
Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
by: Cundy, Chris, et al.
Published: (2020)
by: Cundy, Chris, et al.
Published: (2020)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients
by: Ramjee, Pragnya, et al.
Published: (2024)
by: Ramjee, Pragnya, et al.
Published: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
A summary of the 2023 Society of Obstetric Medicine of Australia and New Zealand (SOMANZ) hypertension in pregnancy guidelines
by: Cathy Latino, et al.
Published: (2024)
by: Cathy Latino, et al.
Published: (2024)
Reassessing the Navier-Stokes Equation and the Yang-Mills Mass Gap Under a Substrate Ontology
by: Zelenka, David D.
Published: (2025)
by: Zelenka, David D.
Published: (2025)
Gaming the Stage: Playable Media and the Rise of English Commercial Theater
by: Bloom, Gina
Published: (2019)
by: Bloom, Gina
Published: (2019)
UK AISI Alignment Evaluation Case-Study
by: Souly, Alexandra, et al.
Published: (2026)
by: Souly, Alexandra, et al.
Published: (2026)
Effects of phenological stages and ensiling length on chemical composition of Megathyrsus maximus ensiled with Moringa oleifera at different proportions
by: Damilola Kola Oyaniran
Published: (2024)
by: Damilola Kola Oyaniran
Published: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
by: Li, Chloe, et al.
Published: (2025)
by: Li, Chloe, et al.
Published: (2025)
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Similar Items
-
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
by: Ayonrinde, Kola
Published: (2024) -
Propensity Inference: Environmental Contributors to LLM Behaviour
by: Järviniemi, Olli, et al.
Published: (2026) -
Progress Measures for Grokking on Real-world Tasks
by: Golechha, Satvik
Published: (2024) -
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
by: Ayonrinde, Kola, et al.
Published: (2025) -
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
by: Ayonrinde, Kola, et al.
Published: (2025)