AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
Fuente:
arXiv
Guardado en:
| Autores principales: | Stix, Charlotte, Pistillo, Matteo, Sastry, Girish, Hobbhahn, Marius, Ortega, Alejandro, Balesni, Mikita, Hallensleben, Annika, Goldowsky-Dill, Nix, Sharkey, Lee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Loss of Control Playbook: Degrees, Dynamics, and Preparedness
por: Stix, Charlotte, et al.
Publicado: (2025)
por: Stix, Charlotte, et al.
Publicado: (2025)
Pre-Deployment Information Sharing: A Zoning Taxonomy for Precursory Capabilities
por: Pistillo, Matteo, et al.
Publicado: (2024)
por: Pistillo, Matteo, et al.
Publicado: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
por: Scheurer, Jérémy, et al.
Publicado: (2023)
por: Scheurer, Jérémy, et al.
Publicado: (2023)
Assurance of Frontier AI Built for National Security
por: Pistillo, Matteo, et al.
Publicado: (2025)
por: Pistillo, Matteo, et al.
Publicado: (2025)
Detecting Strategic Deception Using Linear Probes
por: Goldowsky-Dill, Nicholas, et al.
Publicado: (2025)
por: Goldowsky-Dill, Nicholas, et al.
Publicado: (2025)
Internal Deployment in the AI Act
por: Pistillo, Matteo
Publicado: (2025)
por: Pistillo, Matteo
Publicado: (2025)
Towards evaluations-based safety cases for AI scheming
por: Balesni, Mikita, et al.
Publicado: (2024)
por: Balesni, Mikita, et al.
Publicado: (2024)
Frontier Models are Capable of In-context Scheming
por: Meinke, Alexander, et al.
Publicado: (2024)
por: Meinke, Alexander, et al.
Publicado: (2024)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
por: Braun, Dan, et al.
Publicado: (2024)
por: Braun, Dan, et al.
Publicado: (2024)
Lessons from Studying Two-Hop Latent Reasoning
por: Balesni, Mikita, et al.
Publicado: (2024)
por: Balesni, Mikita, et al.
Publicado: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
por: Laine, Rudolf, et al.
Publicado: (2024)
por: Laine, Rudolf, et al.
Publicado: (2024)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
por: Bushnaq, Lucius, et al.
Publicado: (2024)
por: Bushnaq, Lucius, et al.
Publicado: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
por: Schoen, Bronson, et al.
Publicado: (2025)
por: Schoen, Bronson, et al.
Publicado: (2025)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
por: Bushnaq, Lucius, et al.
Publicado: (2024)
por: Bushnaq, Lucius, et al.
Publicado: (2024)
Children in Police Custody: Adversity and Adversariality Behind Closed Doors
por: Frances Sheahan
Publicado: (2025)
por: Frances Sheahan
Publicado: (2025)
Towards Frontier Safety Policies Plus
por: Pistillo, Matteo
Publicado: (2025)
por: Pistillo, Matteo
Publicado: (2025)
Behind Office Doors
Publicado: (2026)
Publicado: (2026)
Behind Office Doors
Publicado: (2026)
Publicado: (2026)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
por: Bansal, Hritik, et al.
Publicado: (2025)
por: Bansal, Hritik, et al.
Publicado: (2025)
Defending Compute Thresholds Against Legal Loopholes
por: Pistillo, Matteo, et al.
Publicado: (2025)
por: Pistillo, Matteo, et al.
Publicado: (2025)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
por: McKee-Reid, Leo, et al.
Publicado: (2024)
por: McKee-Reid, Leo, et al.
Publicado: (2024)
Chapter 6 The Role of Corporate Governance in Macro-Prudential Regulation of Systemic Risk
por: Dill, Alexander
Publicado: (2020)
por: Dill, Alexander
Publicado: (2020)
Owning the Stuff of Life
por: Stix, Gary
por: Stix, Gary
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
por: Berglund, Lukas, et al.
Publicado: (2023)
por: Berglund, Lukas, et al.
Publicado: (2023)
Behind Closed Doors: An Exploratory Study of the Perceptions of Librarians and the Hidden Intellectual Work of Collection Development in Canadian Public Libraries.
por: Nilsen, Kirsti, et al.
Publicado: (2002)
por: Nilsen, Kirsti, et al.
Publicado: (2002)
Ground states for the Hartree energy functional in the critical case
por: Pistillo, Tommaso
Publicado: (2025)
por: Pistillo, Tommaso
Publicado: (2025)
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
por: Pistillo, Matteo, et al.
Publicado: (2026)
por: Pistillo, Matteo, et al.
Publicado: (2026)
Analogical Reasoning Within a Conceptual Hyperspace
por: Goldowsky, Howard, et al.
Publicado: (2024)
por: Goldowsky, Howard, et al.
Publicado: (2024)
Technical Report: Evaluating Goal Drift in Language Model Agents
por: Arike, Rauno, et al.
Publicado: (2025)
por: Arike, Rauno, et al.
Publicado: (2025)
Forecasting Frontier Language Model Agent Capabilities
por: Pimpale, Govind, et al.
Publicado: (2025)
por: Pimpale, Govind, et al.
Publicado: (2025)
The Day the Library Closed Its Doors
por: Yates, Elizabeth
Publicado: (1970)
por: Yates, Elizabeth
Publicado: (1970)
Hunter Midtown Library: The Closing of an Open Door
por: Foster, Barbara
Publicado: (1976)
por: Foster, Barbara
Publicado: (1976)
DoorBot: Closed-Loop Task Planning and Manipulation for Door Opening in the Wild with Haptic Feedback
por: Wang, Zhi, et al.
Publicado: (2025)
por: Wang, Zhi, et al.
Publicado: (2025)
Bibliophilately Revisited.
por: Nix, Larry T.
Publicado: (2000)
por: Nix, Larry T.
Publicado: (2000)
A Study of the Bookmobile Service of the Madison Public Library.
por: Nix, Larry T.
Publicado: (1981)
por: Nix, Larry T.
Publicado: (1981)
The étale topos reconstructs varieties over sub-p-adic fields
por: Carlson, Magnus, et al.
Publicado: (2024)
por: Carlson, Magnus, et al.
Publicado: (2024)
Monographs in Microform: Issues in Cataloging and Bibliographic Control.
por: Mikita, Elizabeth G.
Publicado: (1981)
por: Mikita, Elizabeth G.
Publicado: (1981)
Large Language Models Often Know When They Are Being Evaluated
por: Needham, Joe, et al.
Publicado: (2025)
por: Needham, Joe, et al.
Publicado: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
por: Højmark, Axel, et al.
Publicado: (2024)
por: Højmark, Axel, et al.
Publicado: (2024)
Ejemplares similares
-
The Loss of Control Playbook: Degrees, Dynamics, and Preparedness
por: Stix, Charlotte, et al.
Publicado: (2025) -
Pre-Deployment Information Sharing: A Zoning Taxonomy for Precursory Capabilities
por: Pistillo, Matteo, et al.
Publicado: (2024) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
por: Scheurer, Jérémy, et al.
Publicado: (2023) -
Assurance of Frontier AI Built for National Security
por: Pistillo, Matteo, et al.
Publicado: (2025) -
Detecting Strategic Deception Using Linear Probes
por: Goldowsky-Dill, Nicholas, et al.
Publicado: (2025)