Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Korbak, Tomek, Balesni, Mikita, Barnes, Elizabeth, Bengio, Yoshua, Benton, Joe, Bloom, Joseph, Chen, Mark, Cooney, Alan, Dafoe, Allan, Dragan, Anca, Emmons, Scott, Evans, Owain, Farhi, David, Greenblatt, Ryan, Hendrycks, Dan, Hobbhahn, Marius, Hubinger, Evan, Irving, Geoffrey, Jenner, Erik, Kokotajlo, Daniel, Krakovna, Victoria, Legg, Shane, Lindner, David, Luan, David, Mądry, Aleksander, Michael, Julian, Nanda, Neel, Orr, Dave, Pachocki, Jakub, Perez, Ethan, Phuong, Mary, Roger, Fabien, Saxe, Joshua, Shlegeris, Buck, Soto, Martín, Steinberger, Eric, Wang, Jasmine, Zaremba, Wojciech, Baker, Bowen, Shah, Rohin, Mikulik, Vlad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
by: Baker, Bowen, et al.
Published: (2025)
by: Baker, Bowen, et al.
Published: (2025)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025)
by: Lindner, David, et al.
Published: (2025)
Evaluating Frontier Models for Stealth and Situational Awareness
by: Phuong, Mary, et al.
Published: (2025)
by: Phuong, Mary, et al.
Published: (2025)
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026)
by: Krakovna, Victoria, et al.
Published: (2026)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
by: Stix, Charlotte, et al.
Published: (2025)
by: Stix, Charlotte, et al.
Published: (2025)
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023)
by: Greenblatt, Ryan, et al.
Published: (2023)
Training Agents to Self-Report Misbehavior
by: Lee, Bruce W., et al.
Published: (2026)
by: Lee, Bruce W., et al.
Published: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
by: Stickland, Asa Cooper, et al.
Published: (2025)
by: Stickland, Asa Cooper, et al.
Published: (2025)
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)
by: Lindner, David, et al.
Published: (2026)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
by: McKee-Reid, Leo, et al.
Published: (2024)
by: McKee-Reid, Leo, et al.
Published: (2024)
Das Berlin Max Webers
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
by: Aldenhoff-Hübinger, Rita, et al.
Published: (2026)
Adsorption design for wastewater treatment / David O. Cooney
by: Cooney, David O
by: Cooney, David O
X-ray reflectivity study of a W/Si multilayer grating
by: P. Mikulík
Published: (2001)
by: P. Mikulík
Published: (2001)
Uncovering Latent Human Wellbeing in Language Model Embeddings
by: Freire, Pedro, et al.
Published: (2024)
by: Freire, Pedro, et al.
Published: (2024)
The Persistent Ambiguity of Adverse Drug Reactions
by: Edward M. Sellers, et al.
Published: (2025)
by: Edward M. Sellers, et al.
Published: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Aligning language models with human preferences
by: Korbak, Tomasz
Published: (2024)
by: Korbak, Tomasz
Published: (2024)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
A Benchmarking Suite for Flexible Job Shop Scheduling Problems with Worker Flexibility under Uncertainty
by: Hutter, David, et al.
Published: (2025)
by: Hutter, David, et al.
Published: (2025)
Frequency Scaling Laws for Flat Plate Wing Active Separation Control
by: Vey, Stefan, et al.
Published: (2025)
by: Vey, Stefan, et al.
Published: (2025)
Safety case template for frontier AI: A cyber inability argument
by: Goemans, Arthur, et al.
Published: (2024)
by: Goemans, Arthur, et al.
Published: (2024)
Red novae, stellar mergers in binary and triple systems, and bipolar nebulae
by: Kaminski, Tomek
Published: (2024)
by: Kaminski, Tomek
Published: (2024)
Emergent Compositional Communication for Latent World Properties
by: Kaszyński, Tomek
Published: (2026)
by: Kaszyński, Tomek
Published: (2026)
In Defiance: 20 Abolitionists You Were Never Taught in Schools. By TomWeiner and AmilcarShabazz. Northampton, MA: Interlink/Olive Branch Press, 2025. 248 pp. $25.00 (paperback). ISBN: 978‐1‐62‐371661‐5
by: Beverly Tomek
Published: (2025)
by: Beverly Tomek
Published: (2025)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Atoms and Molecules as Quantum Attosecond Processors
by: Farhi, Asaf
Published: (2025)
by: Farhi, Asaf
Published: (2025)
On refinements of two-term Machin-like formulas
by: Farhi, Bakir
Published: (2026)
by: Farhi, Bakir
Published: (2026)
Expansion of a bivariate symmetric mean in the neighborhood of the first bisector
by: Farhi, Bakir
Published: (2025)
by: Farhi, Bakir
Published: (2025)
Similar Items
-
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024) -
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025) -
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023) -
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)