Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
Fuente:
arXiv
Saved in:
| Main Authors: | Bengio, Yoshua, Cohen, Michael, Fornasiere, Damiano, Ghosn, Joumana, Greiner, Pietro, MacDermott, Matt, Mindermann, Sören, Oberman, Adam, Richardson, Jesse, Richardson, Oliver, Rondeau, Marc-Antoine, St-Charles, Pierre-Luc, Williams-King, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Language models recognize dropout and Gaussian noise applied to their activations
by: Fornasiere, Damiano, et al.
Published: (2026)
by: Fornasiere, Damiano, et al.
Published: (2026)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Trees and spectra of Heyting algebras
by: Fornasiere, Damiano, et al.
Published: (2024)
by: Fornasiere, Damiano, et al.
Published: (2024)
Intuitionistic Sahlqvist theory for deductive systems
by: Fornasiere, Damiano, et al.
Published: (2022)
by: Fornasiere, Damiano, et al.
Published: (2022)
Whatever Happened to Frank and Fearless?
by: MacDermott, Kathy
Published: (2013)
by: MacDermott, Kathy
Published: (2013)
AI & Human Co-Improvement for Safer Co-Superintelligence
by: Weston, Jason, et al.
Published: (2025)
by: Weston, Jason, et al.
Published: (2025)
Measuring Goal-Directedness
by: MacDermott, Matt, et al.
Published: (2024)
by: MacDermott, Matt, et al.
Published: (2024)
Active Attacks: Red-teaming LLMs via Adaptive Environments
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
by: MacDermott, Matt, et al.
Published: (2025)
by: MacDermott, Matt, et al.
Published: (2025)
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Quantum Mechanics from Symmetry
by: Hegstrom, Roger A., et al.
Published: (2022)
by: Hegstrom, Roger A., et al.
Published: (2022)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
The Reasons that Agents Act: Intention and Instrumental Goals
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
The Alignment Problem from a Deep Learning Perspective
by: Ngo, Richard, et al.
Published: (2022)
by: Ngo, Richard, et al.
Published: (2022)
Bayes-ically fair: A Bayesian Ranking of the Olympic Medal Table
by: MacDermott, Cormac, et al.
Published: (2025)
by: MacDermott, Cormac, et al.
Published: (2025)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Early‐stage evaluation of the Treehouse4Two Retreat Program to support people living with dementia and carer dyads living in rural Victoria, Australia
by: Shahinoor Akter, et al.
Published: (2025)
by: Shahinoor Akter, et al.
Published: (2025)
Local Inconsistency Resolution: The Interplay between Attention and Control in Probabilistic Models
by: Richardson, Oliver E., et al.
Published: (2026)
by: Richardson, Oliver E., et al.
Published: (2026)
Community‐Driven Health Promotion: Evaluation of a Rural Microgrant Program
by: Michele Conlin, et al.
Published: (2024)
by: Michele Conlin, et al.
Published: (2024)
GEORAMA OF TRASH
by: Rania Ghosn
Published: (2016)
by: Rania Ghosn
Published: (2016)
What the ACRL Institute for Information Literacy Best Practices Initiative Tells Us about the Librarian as Teacher.
by: Oberman, Cerise
Published: (2002)
by: Oberman, Cerise
Published: (2002)
Unmasking Technology: A Prelude to Teaching.
by: Oberman, Cerise
Published: (1995)
by: Oberman, Cerise
Published: (1995)
Patterns for Research.
by: Oberman, Cerise
Published: (1984)
by: Oberman, Cerise
Published: (1984)
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
NEPTUNIA PLENA (FABACEAE: MIMOSOIDEAE) REDISCOVERED IN TEXAS
by: Richardson, Alfred, et al.
Published: (2008)
by: Richardson, Alfred, et al.
Published: (2008)
Superintelligence and Law
by: Kolt, Noam
Published: (2026)
by: Kolt, Noam
Published: (2026)
On Generalization for Generative Flow Networks
by: Krichel, Anas, et al.
Published: (2024)
by: Krichel, Anas, et al.
Published: (2024)
Interventional Causal Representation Learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Relative Trajectory Balance is equivalent to Trust-PCL
by: Deleu, Tristan, et al.
Published: (2025)
by: Deleu, Tristan, et al.
Published: (2025)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
by: Yoon, Jaesik, et al.
Published: (2025)
by: Yoon, Jaesik, et al.
Published: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
GFlowNet Foundations
by: Bengio, Yoshua, et al.
Published: (2021)
by: Bengio, Yoshua, et al.
Published: (2021)
Introduction to Bibliography and Research Methods Handbook.
by: Oberman-Soroka, Cerise
Published: (1979)
by: Oberman-Soroka, Cerise
Published: (1979)
Petals Around a Rose: Abstract Reasoning and Bibliographic Instruction.
by: Oberman-Soroka, Cerise
Published: (1980)
by: Oberman-Soroka, Cerise
Published: (1980)
Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
by: Pham, Anh, et al.
Published: (2025)
by: Pham, Anh, et al.
Published: (2025)
Exploring an Allied Health Student Placement Model in Rural Aged‐Care Settings: A UDRH Collaborative Evaluation Using RE‐AIM
by: Carmela Leone, et al.
Published: (2026)
by: Carmela Leone, et al.
Published: (2026)
Does A Dietitian‐Led Celiac Disease Clinic (DLCC) Facilitate Timely Diagnosis and Nutrition Care for Patients With Celiac Disease?
by: M. Palmer, et al.
Published: (2025)
by: M. Palmer, et al.
Published: (2025)
Similar Items
-
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024) -
Language models recognize dropout and Gaussian noise applied to their activations
by: Fornasiere, Damiano, et al.
Published: (2026) -
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025) -
Trees and spectra of Heyting algebras
by: Fornasiere, Damiano, et al.
Published: (2024) -
Intuitionistic Sahlqvist theory for deductive systems
by: Fornasiere, Damiano, et al.
Published: (2022)