Auditing Games for Sandbagging
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taylor, Jordan, Black, Sid, Bowen, Dillon, Read, Thomas, Golechha, Satvik, Zelenka-Martin, Alex, Makins, Oliver, Kissane, Connor, Ayonrinde, Kola, Merizian, Jacob, Marks, Samuel, Cundy, Chris, Bloom, Joseph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024)
von: Ayonrinde, Kola
Veröffentlicht: (2024)
Propensity Inference: Environmental Contributors to LLM Behaviour
von: Järviniemi, Olli, et al.
Veröffentlicht: (2026)
von: Järviniemi, Olli, et al.
Veröffentlicht: (2026)
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024)
von: Golechha, Satvik
Veröffentlicht: (2024)
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
von: Dombrowski, Ann-Kathrin, et al.
Veröffentlicht: (2025)
von: Dombrowski, Ann-Kathrin, et al.
Veröffentlicht: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
von: Chanin, David, et al.
Veröffentlicht: (2024)
von: Chanin, David, et al.
Veröffentlicht: (2024)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
From Mechanistic to Compositional Interpretability
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
von: Gauderis, Ward, et al.
Veröffentlicht: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
Studying Cross-cluster Modularity in Neural Networks
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
von: Cundy, Chris, et al.
Veröffentlicht: (2023)
von: Cundy, Chris, et al.
Veröffentlicht: (2023)
More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration
von: Yadav, Advait, et al.
Veröffentlicht: (2026)
von: Yadav, Advait, et al.
Veröffentlicht: (2026)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
Do Large Language Models Know What They Are Capable Of?
von: Barkan, Casey O., et al.
Veröffentlicht: (2025)
von: Barkan, Casey O., et al.
Veröffentlicht: (2025)
Sandbagging in a Simple Survival Bandit Problem
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
von: Dyer, Joel, et al.
Veröffentlicht: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
von: Ryd, Emil, et al.
Veröffentlicht: (2026)
von: Ryd, Emil, et al.
Veröffentlicht: (2026)
Algunas observaciones sobre los regímenes especiales de seguridad social en la agricultura
von: A. Zelenka
Veröffentlicht: (1963)
von: A. Zelenka
Veröffentlicht: (1963)
Some remarks on special social security schemes for agriculture
von: A. Zelenka
Veröffentlicht: (1963)
von: A. Zelenka
Veröffentlicht: (1963)
Quelques remarques sur les régimes spéciaux de sécurité sociale des agriculteurs
von: A. Zelenka
Veröffentlicht: (1963)
von: A. Zelenka
Veröffentlicht: (1963)
Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
von: Cundy, Chris, et al.
Veröffentlicht: (2020)
von: Cundy, Chris, et al.
Veröffentlicht: (2020)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
CataractBot: An LLM-Powered Expert-in-the-Loop Chatbot for Cataract Patients
von: Ramjee, Pragnya, et al.
Veröffentlicht: (2024)
von: Ramjee, Pragnya, et al.
Veröffentlicht: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
A summary of the 2023 Society of Obstetric Medicine of Australia and New Zealand (SOMANZ) hypertension in pregnancy guidelines
von: Cathy Latino, et al.
Veröffentlicht: (2024)
von: Cathy Latino, et al.
Veröffentlicht: (2024)
Reassessing the Navier-Stokes Equation and the Yang-Mills Mass Gap Under a Substrate Ontology
von: Zelenka, David D.
Veröffentlicht: (2025)
von: Zelenka, David D.
Veröffentlicht: (2025)
Gaming the Stage: Playable Media and the Rise of English Commercial Theater
von: Bloom, Gina
Veröffentlicht: (2019)
von: Bloom, Gina
Veröffentlicht: (2019)
UK AISI Alignment Evaluation Case-Study
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
von: Souly, Alexandra, et al.
Veröffentlicht: (2026)
Effects of phenological stages and ensiling length on chemical composition of Megathyrsus maximus ensiled with Moringa oleifera at different proportions
von: Damilola Kola Oyaniran
Veröffentlicht: (2024)
von: Damilola Kola Oyaniran
Veröffentlicht: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
von: Li, Chloe, et al.
Veröffentlicht: (2025)
von: Li, Chloe, et al.
Veröffentlicht: (2025)
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Ähnliche Einträge
-
Adaptive Sparse Allocation with Mutual Choice & Feature Choice Sparse Autoencoders
von: Ayonrinde, Kola
Veröffentlicht: (2024) -
Propensity Inference: Environmental Contributors to LLM Behaviour
von: Järviniemi, Olli, et al.
Veröffentlicht: (2026) -
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024) -
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025) -
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
von: Ayonrinde, Kola, et al.
Veröffentlicht: (2025)