Alignment faking in large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Greenblatt, Ryan, Denison, Carson, Wright, Benjamin, Roger, Fabien, MacDiarmid, Monte, Marks, Sam, Treutlein, Johannes, Belonax, Tim, Chen, Jack, Duvenaud, David, Khan, Akbir, Michael, Julian, Mindermann, Sören, Perez, Ethan, Petrini, Linda, Uesato, Jonathan, Kaplan, Jared, Shlegeris, Buck, Bowman, Samuel R., Hubinger, Evan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023)
by: Greenblatt, Ryan, et al.
Published: (2023)
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025)
by: MacDiarmid, Monte, et al.
Published: (2025)
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024)
by: Benton, Joe, et al.
Published: (2024)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Auditing language models for hidden objectives
by: Marks, Samuel, et al.
Published: (2025)
by: Marks, Samuel, et al.
Published: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
by: Lynch, Aengus, et al.
Published: (2025)
by: Lynch, Aengus, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
A Post‐Pandemic Bail System: Lessons Learned From Supervising Accused During Covid‐19
by: Laura MacDiarmid, et al.
Published: (2025)
by: Laura MacDiarmid, et al.
Published: (2025)
Scottish meat consumption survey (Feb–Jul 2023): attitudes, COM-B measures, and perceived effectiveness of meat-reduction policies (Best-Worst Scaling)
by: McBey, David, et al.
Published: (2026)
by: McBey, David, et al.
Published: (2026)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
Ctrl-Z: Controlling AI Agents via Resampling
by: Bhatt, Aryan, et al.
Published: (2025)
by: Bhatt, Aryan, et al.
Published: (2025)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
by: Järviniemi, Olli, et al.
Published: (2024)
by: Järviniemi, Olli, et al.
Published: (2024)
Steering Language Models With Activation Engineering
by: Turner, Alexander Matt, et al.
Published: (2023)
by: Turner, Alexander Matt, et al.
Published: (2023)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
The Alignment Problem from a Deep Learning Perspective
by: Ngo, Richard, et al.
Published: (2022)
by: Ngo, Richard, et al.
Published: (2022)
Das Berlin Max Webers
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026)
by: Gan, Eric, et al.
Published: (2026)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
by: Maiya, Sharan, et al.
Published: (2025)
by: Maiya, Sharan, et al.
Published: (2025)
Stress-Testing Capability Elicitation With Password-Locked Models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Intraoperative Paragastric Vagal Nerve Block With Ropivacaine Reduces Postoperative Nausea and Pain on the Day of Surgery After Laparoscopic Sleeve Gastrectomy: A Retrospective Cohort Study With Propensity Score Matching
by: Susumu Inamine, et al.
Published: (2026)
by: Susumu Inamine, et al.
Published: (2026)
Hessian determinants and averaging operators over surfaces in ${\mathbb R}^3$
by: Greenblatt, Michael
Published: (2021)
by: Greenblatt, Michael
Published: (2021)
Oscillatory integrals and weighted gradient flows
by: Greenblatt, Michael
Published: (2024)
by: Greenblatt, Michael
Published: (2024)
Convexity, Fourier transforms, and lattice point discrepancy
by: Greenblatt, Michael
Published: (2024)
by: Greenblatt, Michael
Published: (2024)
A method for bounding oscillatory integrals in terms of non-oscillatory integrals
by: Greenblatt, Michael
Published: (2022)
by: Greenblatt, Michael
Published: (2022)
Expanding Children's Programming in School and Public Libraries.
by: Greenblatt, Melinda
Published: (1979)
by: Greenblatt, Melinda
Published: (1979)
Pendant appearances and components in random graphs from structured classes
by: McDiarmid, Colin
Published: (2021)
by: McDiarmid, Colin
Published: (2021)
The Savage Worlds of Henry Drummond (1851–1897): Science, Racism and Religion in the Work of a Popular Evolutionist
by: Diarmid A. Finnegan
Published: (2025)
by: Diarmid A. Finnegan
Published: (2025)
Die Rolle gemeinwohlorientierter Akteure zur Unterstützung gemeinschaftlicher Wohnprojekte in Berlin
by: Hübinger, Heike, et al.
Published: (2022)
by: Hübinger, Heike, et al.
Published: (2022)
Language Models Learn to Mislead Humans via RLHF
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
Explaining Neural Scaling Laws
by: Bahri, Yasaman, et al.
Published: (2021)
by: Bahri, Yasaman, et al.
Published: (2021)
The boundary disorder correlation for the Ising model on a cylinder
by: Greenblatt, Rafael Leon
Published: (2024)
by: Greenblatt, Rafael Leon
Published: (2024)
Constructing a weakly-interacting fixed point of the Fermionic Polchinski equation
by: Greenblatt, Rafael Leon
Published: (2024)
by: Greenblatt, Rafael Leon
Published: (2024)
Similar Items
-
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024) -
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023) -
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025) -
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024) -
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)