Sabotage Evaluations for Frontier Models
Fuente:
arXiv
Saved in:
| Main Authors: | Benton, Joe, Wagner, Misha, Christiansen, Eric, Anil, Cem, Perez, Ethan, Srivastav, Jai, Durmus, Esin, Ganguli, Deep, Kravec, Shauna, Shlegeris, Buck, Kaplan, Jared, Karnofsky, Holden, Hubinger, Evan, Grosse, Roger, Bowman, Samuel R., Duvenaud, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026)
by: Gan, Eric, et al.
Published: (2026)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024)
by: Huang, Saffron, et al.
Published: (2024)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Towards Understanding Sycophancy in Language Models
by: Sharma, Mrinank, et al.
Published: (2023)
by: Sharma, Mrinank, et al.
Published: (2023)
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025)
by: Kutasov, Jon, et al.
Published: (2025)
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
by: Handa, Kunal, et al.
Published: (2025)
by: Handa, Kunal, et al.
Published: (2025)
Das Berlin Max Webers
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023)
by: Greenblatt, Ryan, et al.
Published: (2023)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Sabotage the Mantel Theorem
by: Behague, Natalie, et al.
Published: (2025)
by: Behague, Natalie, et al.
Published: (2025)
Quantum Sabotage Complexity
by: Cornelissen, Arjan, et al.
Published: (2024)
by: Cornelissen, Arjan, et al.
Published: (2024)
Towards Measuring the Representation of Subjective Global Opinions in Language Models
by: Durmus, Esin, et al.
Published: (2023)
by: Durmus, Esin, et al.
Published: (2023)
Clio: Privacy-Preserving Insights into Real-World AI Use
by: Tamkin, Alex, et al.
Published: (2024)
by: Tamkin, Alex, et al.
Published: (2024)
Complete Game Logic with Sabotage
by: Wafa, Noah Abou El, et al.
Published: (2024)
by: Wafa, Noah Abou El, et al.
Published: (2024)
Die Rolle gemeinwohlorientierter Akteure zur Unterstützung gemeinschaftlicher Wohnprojekte in Berlin
by: Hübinger, Heike, et al.
Published: (2022)
by: Hübinger, Heike, et al.
Published: (2022)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
by: Järviniemi, Olli, et al.
Published: (2024)
by: Järviniemi, Olli, et al.
Published: (2024)
Teaching Physics with Sabotage and SimShield
by: Rosehill, Daniel, et al.
Published: (2026)
by: Rosehill, Daniel, et al.
Published: (2026)
How Government Experts Self-Sabotage
by: Gerblinger, Christiane
Published: (2023)
by: Gerblinger, Christiane
Published: (2023)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
by: Gringras, David, et al.
Published: (2026)
by: Gringras, David, et al.
Published: (2026)
“One Is a Frontier”: Settler Migration as Transmogrification
by: Joseph Kaplan Weinger
Published: (2025)
by: Joseph Kaplan Weinger
Published: (2025)
Strategies in Sabotage Games: Temporal and Epistemic Perspectives
by: Gierasimczuk, Nina, et al.
Published: (2026)
by: Gierasimczuk, Nina, et al.
Published: (2026)
Silent Sabotage: A Challenge in Cardiac Resynchronization
by: Mariana Pereira Santos, et al.
Published: (2026)
by: Mariana Pereira Santos, et al.
Published: (2026)
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
by: Gligoric, Kristina, et al.
Published: (2024)
by: Gligoric, Kristina, et al.
Published: (2024)
Stress-Testing Model Specs Reveals Character Differences among Language Models
by: Zhang, Jifan, et al.
Published: (2025)
by: Zhang, Jifan, et al.
Published: (2025)
Elektrosynthese mit Sonne und Wind
by: Emil Roduner, et al.
Published: (2025)
by: Emil Roduner, et al.
Published: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
HPLC‐MS/MS Method for Monitoring of L‐Asparaginase Activity by Using Asparagine and Aspartic Acid Plasma Levels
by: Cem Kaplan, et al.
Published: (2026)
by: Cem Kaplan, et al.
Published: (2026)
Sojourner under Sabotage: A Serious Testing and Debugging Game
by: Straubinger, Philipp, et al.
Published: (2025)
by: Straubinger, Philipp, et al.
Published: (2025)
Ctrl-Z: Controlling AI Agents via Resampling
by: Bhatt, Aryan, et al.
Published: (2025)
by: Bhatt, Aryan, et al.
Published: (2025)
Doppler Ultrasonography in the Diagnosis of Dogs With Benign Prostatic Hyperplasia and Investigation of the Efficacy of Finasteride in Treatment
by: Burcu Esin, et al.
Published: (2025)
by: Burcu Esin, et al.
Published: (2025)
Similar Items
-
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024) -
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026) -
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025) -
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)