Saved in:
| Main Authors: | Phuong, Mary, Aitchison, Matthew, Catt, Elliot, Cogan, Sarah, Kaskasoli, Alexandre, Krakovna, Victoria, Lindner, David, Rahtz, Matthew, Assael, Yannis, Hodkinson, Sarah, Howard, Heidi, Lieberum, Tom, Kumar, Ramana, Raad, Maria Abi, Webson, Albert, Ho, Lewis, Lin, Sharon, Farquhar, Sebastian, Hutter, Marcus, Deletang, Gregoire, Ruoss, Anian, El-Sayed, Seliem, Brown, Sasha, Dragan, Anca, Shah, Rohin, Dafoe, Allan, Shevlane, Toby |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.13793 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributional Bellman Operators over Mean Embeddings
by: Wenliang, Li Kevin, et al.
Published: (2023)
by: Wenliang, Li Kevin, et al.
Published: (2023)
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024)
by: Grau-Moya, Jordi, et al.
Published: (2024)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
Evaluating Frontier Models for Stealth and Situational Awareness
by: Phuong, Mary, et al.
Published: (2025)
by: Phuong, Mary, et al.
Published: (2025)
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026)
by: Krakovna, Victoria, et al.
Published: (2026)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)
by: Lindner, David, et al.
Published: (2026)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024)
by: Heurtel-Depeiges, David, et al.
Published: (2024)
Why is prompting hard? Understanding prompts on binary sequence predictors
by: Wenliang, Li Kevin, et al.
Published: (2025)
by: Wenliang, Li Kevin, et al.
Published: (2025)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
by: Kramár, János, et al.
Published: (2024)
by: Kramár, János, et al.
Published: (2024)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025)
by: Genewein, Tim, et al.
Published: (2025)
Estandarización educativa en Chile: tensiones y consecuencias para el trabajo docente
by: Jenny Assaél
Published: (2018)
by: Jenny Assaél
Published: (2018)
Les Mysteres de L’onomastique Dans L’helene D’euripide
by: Jacqueline Assaël
Published: (2014)
by: Jacqueline Assaël
Published: (2014)
El pensamiento de la CEPAL : un intento de evaluar algunas críticas a sus ideas principales / Héctor Assael
by: Assael, Héctor
Published: (1981)
by: Assael, Héctor
Published: (1981)
El mito del subterráneo: memoria, política y participación en un liceo secundario de Santiago
by: Jenny Assaél
Published: (2001)
by: Jenny Assaél
Published: (2001)
Auf Pump
by: Ruoss, Matthias
Published: (2025)
by: Ruoss, Matthias
Published: (2025)
Schweizerdeutsch und Sprachbewusstsein
by: Ruoss, Emanuel
Published: (2020)
by: Ruoss, Emanuel
Published: (2020)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
by: El-Sayed, Seliem, et al.
Published: (2024)
by: El-Sayed, Seliem, et al.
Published: (2024)
La empresa educativa Chilena
by: Jenny Assaél Budnik
Published: (2011)
by: Jenny Assaél Budnik
Published: (2011)
Acerca del lugar de la entrevista en encuestas electorales
by: Assael Ortiz Lazcano
Published: (2006)
by: Assael Ortiz Lazcano
Published: (2006)
Masculinity and Danger on the Eighteenth-Century Grand Tour
by: Goldsmith, Sarah
Published: (2021)
by: Goldsmith, Sarah
Published: (2021)
A Dangerous Occupation? Violence in Public Libraries.
by: Farrugia, Sarah
Published: (2002)
by: Farrugia, Sarah
Published: (2002)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
An Information-Theoretic Diagnostic Analytics Framework for Mapping Past-Future Dependence in Horizon-Specific Forecastability
by: Catt, Peter Maurice
Published: (2026)
by: Catt, Peter Maurice
Published: (2026)
Forecastability as an Information-Theoretic Limit on Prediction
by: Catt, Peter Maurice
Published: (2026)
by: Catt, Peter Maurice
Published: (2026)
On the Limits of Prediction: Forecastability Profiles and Information Decay in Time Series
by: Catt, Peter Maurice
Published: (2026)
by: Catt, Peter Maurice
Published: (2026)
The Olympic Training Field for Planning Quality Library Services.
by: Catt, Martha E.
Published: (1995)
by: Catt, Martha E.
Published: (1995)
Navigating the Political Dangers of Critical Approaches to Sexual Violence Work
by: Sarah Socorro Hurtado
Published: (2024)
by: Sarah Socorro Hurtado
Published: (2024)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025)
by: Farquhar, Sebastian, et al.
Published: (2025)
Comprehensive AI governance requires addressing non-model gains
by: Goemans, Arthur, et al.
Published: (2026)
by: Goemans, Arthur, et al.
Published: (2026)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Canonicity in power and modal logics of finite achronal width
by: Goldblatt, Robert, et al.
Published: (2022)
by: Goldblatt, Robert, et al.
Published: (2022)
Groups with ET0L co-word problem
by: Kohli, Raad Al, et al.
Published: (2024)
by: Kohli, Raad Al, et al.
Published: (2024)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Flow over inclined flat plate
by: Kappillil Rajeev, Rohin
Published: (2025)
by: Kappillil Rajeev, Rohin
Published: (2025)
Quantifying stability of non-power-seeking in artificial agents
by: Gunter, Evan Ryan, et al.
Published: (2024)
by: Gunter, Evan Ryan, et al.
Published: (2024)
Academic drifts in vocational, professional, and continuing education: A multi-perspective approach for the case of Switzerland
by: Neumann, Jörg, et al.
Published: (2025)
by: Neumann, Jörg, et al.
Published: (2025)
Mastering Board Games by External and Internal Planning with Language Models
by: Schultz, John, et al.
Published: (2024)
by: Schultz, John, et al.
Published: (2024)
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
by: Mueller, Aaron, et al.
Published: (2023)
by: Mueller, Aaron, et al.
Published: (2023)
Similar Items
-
Distributional Bellman Operators over Mean Embeddings
by: Wenliang, Li Kevin, et al.
Published: (2023) -
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024) -
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023) -
Evaluating Frontier Models for Stealth and Situational Awareness
by: Phuong, Mary, et al.
Published: (2025) -
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026)