Lessons from External Review of DeepMind's Scheming Inability Safety Case
Fuente:
arXiv
Saved in:
| Main Authors: | Barrett, Stephen, Zabala, Francisco Javier Campos, Fillingham, Sean P., Siddique, Umair, Walpole, James, Bloomfield, Robin, Papadatos, Henry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Competence Shadow: Theory and Bounds of AI Assistance in Safety Engineering
by: Siddique, Umair
Published: (2026)
by: Siddique, Umair
Published: (2026)
Reinforcement Learning in Strategy-Based and Atari Games: A Review of Google DeepMinds Innovations
by: Shaheen, Abdelrhman, et al.
Published: (2025)
by: Shaheen, Abdelrhman, et al.
Published: (2025)
Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent
by: Hoppmann, Björn, et al.
Published: (2026)
by: Hoppmann, Björn, et al.
Published: (2026)
Adaptation of Google DeepMind AI AlphaFold2 in Single Protein 3D Structure Prediction: A Critical Review
by: Madawalagama, Erandi Prabashani
Published: (2024)
by: Madawalagama, Erandi Prabashani
Published: (2024)
Evaluating AI Providers' Frontier Safety Frameworks
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
Linear Probe Penalties Reduce LLM Sycophancy
by: Papadatos, Henry, et al.
Published: (2024)
by: Papadatos, Henry, et al.
Published: (2024)
Confidence in Assurance 2.0 Cases
by: Bloomfield, Robin, et al.
Published: (2024)
by: Bloomfield, Robin, et al.
Published: (2024)
Designing Reliable Experiments with Generative Agent-Based Modeling: A Comprehensive Guide Using Concordia by Google DeepMind
by: Navarro, Alejandro Leonardo García, et al.
Published: (2024)
by: Navarro, Alejandro Leonardo García, et al.
Published: (2024)
A Logic of Inability
by: Wang, Shanxia
Published: (2026)
by: Wang, Shanxia
Published: (2026)
STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems
by: Barrett, Steve, et al.
Published: (2025)
by: Barrett, Steve, et al.
Published: (2025)
Understanding: reframing automation and assurance
by: Bloomfield, Robin
Published: (2026)
by: Bloomfield, Robin
Published: (2026)
Publications of the Office of Education, 1963. Bulletin, 1963, No. 29. OE-11000C
by: Walpole, Martha
Published: (1963)
by: Walpole, Martha
Published: (1963)
Optimisation of Pulse Waveforms for Qubit Gates using Deep Learning
by: Fillingham, Zachary, et al.
Published: (2024)
by: Fillingham, Zachary, et al.
Published: (2024)
Probabilidad y estadística / Ronald E. Walpole, Raymond H. Myers ; traducción de Gerardo Maldonado V zquez
by: Walpole, Ronald E
Published: (1992)
by: Walpole, Ronald E
Published: (1992)
Under Construction: Identity and Isomorphism in the Merger of a Library and Information Science School and an Education School.
by: Walpole, MaryBeth
Published: (2000)
by: Walpole, MaryBeth
Published: (2000)
A Methodology for Quantitative AI Risk Modeling
by: Murray, Malcolm, et al.
Published: (2025)
by: Murray, Malcolm, et al.
Published: (2025)
Assessing Confidence with Assurance 2.0
by: Bloomfield, Robin, et al.
Published: (2022)
by: Bloomfield, Robin, et al.
Published: (2022)
Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems
by: Bloomfield, Robin, et al.
Published: (2025)
by: Bloomfield, Robin, et al.
Published: (2025)
Assurance of AI Systems From a Dependability Perspective
by: Bloomfield, Robin, et al.
Published: (2024)
by: Bloomfield, Robin, et al.
Published: (2024)
Quantifying Confidence in Assurance 2.0 Arguments
by: Bloomfield, Robin, et al.
Published: (2026)
by: Bloomfield, Robin, et al.
Published: (2026)
Mapping AI Benchmark Data to Quantitative Risk Estimates Through Expert Elicitation
by: Murray, Malcolm, et al.
Published: (2025)
by: Murray, Malcolm, et al.
Published: (2025)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
by: Dassanayake, Rishane, et al.
Published: (2025)
by: Dassanayake, Rishane, et al.
Published: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Entangled states dynamics of moving two-level atoms in a thermal field bath
by: Papadatos, Nikolaos, et al.
Published: (2023)
by: Papadatos, Nikolaos, et al.
Published: (2023)
2023 International Literacy Association's Timothy and Cynthia Shanahan Outstanding Dissertation Award
by: Sharon Walpole, et al.
Published: (2024)
by: Sharon Walpole, et al.
Published: (2024)
2024 International Literacy Association's Timothy and Cynthia Shanahan Outstanding Dissertation Award
by: Sharon Walpole, et al.
Published: (2024)
by: Sharon Walpole, et al.
Published: (2024)
A Digital Twin Framework for Metamorphic Testing of Autonomous Driving Systems Using Generative Model
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
AI-Augmented Metamorphic Testing for Comprehensive Validation of Autonomous Vehicles
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
by: Campos, Simeon, et al.
Published: (2025)
by: Campos, Simeon, et al.
Published: (2025)
The Role of Risk Modeling in Advanced AI Risk Management
by: Touzet, Chloé, et al.
Published: (2025)
by: Touzet, Chloé, et al.
Published: (2025)
Free Falling: Fizzing Wild‐Caught Smallmouth Bass Results in the Inability to Control Buoyancy in Deep Water
by: Jamie C. Madden, et al.
Published: (2024)
by: Jamie C. Madden, et al.
Published: (2024)
Inability of spatial transformations of CNN feature maps to support invariant recognition
by: Jansson, Ylva, et al.
Published: (2020)
by: Jansson, Ylva, et al.
Published: (2020)
Beyond Ability: The Four-Fold Spectrum of Power and the Logic of Full Inability
by: Wang, Shanxia
Published: (2026)
by: Wang, Shanxia
Published: (2026)
Regaining Gait Independence in Idiopathic Inflammatory Myopathy Patients With Inability to Ambulate
by: Naoki Mugii, et al.
Published: (2025)
by: Naoki Mugii, et al.
Published: (2025)
Defeaters and Eliminative Argumentation in Assurance 2.0
by: Bloomfield, Robin, et al.
Published: (2024)
by: Bloomfield, Robin, et al.
Published: (2024)
The Quantum Mechanics of Minds and Worlds
by: A. Barrett, Jeffrey
Published: (2026)
by: A. Barrett, Jeffrey
Published: (2026)
Evaluating the Goal-Directedness of Large Language Models
by: Everitt, Tom, et al.
Published: (2025)
by: Everitt, Tom, et al.
Published: (2025)
The importance of gas starvation in driving satellite quenching in galaxy groups at $z\sim 0.8$
by: Baxter, Devontae C., et al.
Published: (2024)
by: Baxter, Devontae C., et al.
Published: (2024)
Die Sprache
by: Bloomfield, Leonard
Published: (2014)
by: Bloomfield, Leonard
Published: (2014)
How everything works: making physics out of the ordinary / Louis A. Bloomfield
by: Bloomfield, Louis
Published: (2007)
by: Bloomfield, Louis
Published: (2007)
Similar Items
-
The Competence Shadow: Theory and Bounds of AI Assistance in Safety Engineering
by: Siddique, Umair
Published: (2026) -
Reinforcement Learning in Strategy-Based and Atari Games: A Review of Google DeepMinds Innovations
by: Shaheen, Abdelrhman, et al.
Published: (2025) -
Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent
by: Hoppmann, Björn, et al.
Published: (2026) -
Adaptation of Google DeepMind AI AlphaFold2 in Single Protein 3D Structure Prediction: A Critical Review
by: Madawalagama, Erandi Prabashani
Published: (2024) -
Evaluating AI Providers' Frontier Safety Frameworks
by: Stelling, Lily, et al.
Published: (2025)