An Approach to Technical AGI Safety and Security
Fuente:
arXiv
Salvato in:
| Autori principali: | Shah, Rohin, Irpan, Alex, Turner, Alexander Matt, Wang, Anna, Conmy, Arthur, Lindner, David, Brown-Cohen, Jonah, Ho, Lewis, Nanda, Neel, Popa, Raluca Ada, Jain, Rishub, Greig, Rory, Albanie, Samuel, Emmons, Scott, Farquhar, Sebastian, Krier, Sébastien, Rajamanoharan, Senthooran, Bridgers, Sophie, Ijitoye, Tobi, Everitt, Tom, Krakovna, Victoria, Varma, Vikrant, Mikulik, Vladimir, Kenton, Zachary, Orr, Dave, Legg, Shane, Goodman, Noah, Dafoe, Allan, Flynn, Four, Dragan, Anca |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Human-AI Complementarity: A Goal for Amplified Oversight
di: Jain, Rishub, et al.
Pubblicazione: (2025)
di: Jain, Rishub, et al.
Pubblicazione: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
di: Lieberum, Tom, et al.
Pubblicazione: (2024)
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024)
Subliminal Learning Is Steering Vector Distillation
di: Blank, Camila, et al.
Pubblicazione: (2026)
di: Blank, Camila, et al.
Pubblicazione: (2026)
Realistic honeypot evaluations for scheming propensity
di: Krakovna, Victoria, et al.
Pubblicazione: (2026)
di: Krakovna, Victoria, et al.
Pubblicazione: (2026)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
di: Emmons, Scott, et al.
Pubblicazione: (2025)
di: Emmons, Scott, et al.
Pubblicazione: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
Levels of AGI for Operationalizing Progress on the Path to AGI
di: Morris, Meredith Ringel, et al.
Pubblicazione: (2023)
di: Morris, Meredith Ringel, et al.
Pubblicazione: (2023)
How Well Do Models Follow Their Constitutions?
di: Jakkli, Arya, et al.
Pubblicazione: (2026)
di: Jakkli, Arya, et al.
Pubblicazione: (2026)
Eliciting Secret Knowledge from Language Models
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
di: Rodriguez, Mikel, et al.
Pubblicazione: (2025)
di: Rodriguez, Mikel, et al.
Pubblicazione: (2025)
Gram: Assessing sabotage propensities via automated alignment auditing
di: Lindner, David, et al.
Pubblicazione: (2026)
di: Lindner, David, et al.
Pubblicazione: (2026)
AGI, Governments, and Free Societies
di: Bullock, Justin B., et al.
Pubblicazione: (2025)
di: Bullock, Justin B., et al.
Pubblicazione: (2025)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
di: Cywiński, Bartosz, et al.
Pubblicazione: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
di: Soligo, Anna, et al.
Pubblicazione: (2026)
di: Soligo, Anna, et al.
Pubblicazione: (2026)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
di: Ferrando, Javier, et al.
Pubblicazione: (2024)
di: Ferrando, Javier, et al.
Pubblicazione: (2024)
Convergent Linear Representations of Emergent Misalignment
di: Soligo, Anna, et al.
Pubblicazione: (2025)
di: Soligo, Anna, et al.
Pubblicazione: (2025)
Evaluating Frontier Models for Stealth and Situational Awareness
di: Phuong, Mary, et al.
Pubblicazione: (2025)
di: Phuong, Mary, et al.
Pubblicazione: (2025)
Distributional AGI Safety
di: Tomašev, Nenad, et al.
Pubblicazione: (2025)
di: Tomašev, Nenad, et al.
Pubblicazione: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
di: Macar, Uzay, et al.
Pubblicazione: (2025)
di: Macar, Uzay, et al.
Pubblicazione: (2025)
Model Organisms for Emergent Misalignment
di: Turner, Edward, et al.
Pubblicazione: (2025)
di: Turner, Edward, et al.
Pubblicazione: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
di: Wang, Atticus, et al.
Pubblicazione: (2025)
di: Wang, Atticus, et al.
Pubblicazione: (2025)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
di: Emmons, Scott, et al.
Pubblicazione: (2025)
di: Emmons, Scott, et al.
Pubblicazione: (2025)
Measuring Progress Toward AGI: A Cognitive Framework
di: Burnell, Ryan, et al.
Pubblicazione: (2026)
di: Burnell, Ryan, et al.
Pubblicazione: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
Consistency Training Helps Stop Sycophancy and Jailbreaks
di: Irpan, Alex, et al.
Pubblicazione: (2025)
di: Irpan, Alex, et al.
Pubblicazione: (2025)
X-ray reflectivity study of a W/Si multilayer grating
di: P. Mikulík
Pubblicazione: (2001)
di: P. Mikulík
Pubblicazione: (2001)
A Rosetta Stone for AI Benchmarks
di: Ho, Anson, et al.
Pubblicazione: (2025)
di: Ho, Anson, et al.
Pubblicazione: (2025)
Can AI mediation improve democratic deliberation?
di: Tessler, Michael Henry, et al.
Pubblicazione: (2026)
di: Tessler, Michael Henry, et al.
Pubblicazione: (2026)
EXPLORING THE MECHANISM OF THE ELECTRONIC QUENCHING OF NO(A2Σ+) WITH CH4, CH3OH, CO2, AND C2H2
di: Bridgers-Anguiano, Aerial
Pubblicazione: (2024)
di: Bridgers-Anguiano, Aerial
Pubblicazione: (2024)
Dense SAE Latents Are Features, Not Bugs
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
Incentives for Responsiveness, Instrumental Control and Impact
di: Carey, Ryan, et al.
Pubblicazione: (2020)
di: Carey, Ryan, et al.
Pubblicazione: (2020)
Restoring Relational Balance: Family Therapy Through the CAT ‐ FAWN Indigenous Lens
di: Don Four Arrows Jacobs
Pubblicazione: (2026)
di: Don Four Arrows Jacobs
Pubblicazione: (2026)
Adequate Support of Public Libraries
di: Conmy, Peter T.
Pubblicazione: (1978)
di: Conmy, Peter T.
Pubblicazione: (1978)
The Public Library in Transition.
di: Conmy, Peter T.
Pubblicazione: (1982)
di: Conmy, Peter T.
Pubblicazione: (1982)
The Public Library and the Family.
di: Conmy, Peter T.
Pubblicazione: (1980)
di: Conmy, Peter T.
Pubblicazione: (1980)
Building Production-Ready Probes For Gemini
di: Kramár, János, et al.
Pubblicazione: (2026)
di: Kramár, János, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Human-AI Complementarity: A Goal for Amplified Oversight
di: Jain, Rishub, et al.
Pubblicazione: (2025) -
Improving Dictionary Learning with Gated Sparse Autoencoders
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024) -
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
di: Lieberum, Tom, et al.
Pubblicazione: (2024) -
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
di: Rajamanoharan, Senthooran, et al.
Pubblicazione: (2024) -
Subliminal Learning Is Steering Vector Distillation
di: Blank, Camila, et al.
Pubblicazione: (2026)