Frontier Models are Capable of In-context Scheming
Fuente:
arXiv
Saved in:
| Main Authors: | Meinke, Alexander, Schoen, Bronson, Scheurer, Jérémy, Balesni, Mikita, Shah, Rusheb, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
by: McKee-Reid, Leo, et al.
Published: (2024)
by: McKee-Reid, Leo, et al.
Published: (2024)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
TracrBench: Generating Interpretability Testbeds with Large Language Models
by: Thurnherr, Hannes, et al.
Published: (2024)
by: Thurnherr, Hannes, et al.
Published: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Jailbroken Frontier Models Retain Their Capabilities
by: Zhu, Daniel, et al.
Published: (2026)
by: Zhu, Daniel, et al.
Published: (2026)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Masked Generative Priors Improve World Models Sequence Modelling Capabilities
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving
by: Abouelazm, Ahmed, et al.
Published: (2024)
by: Abouelazm, Ahmed, et al.
Published: (2024)
Early Signs of Steganographic Capabilities in Frontier LLMs
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Training Language Models with Language Feedback at Scale
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Learning of Population Dynamics: Inverse Optimization Meets JKO Scheme
by: Persiianov, Mikhail, et al.
Published: (2025)
by: Persiianov, Mikhail, et al.
Published: (2025)
Compute Requirements for Algorithmic Innovation in Frontier AI Models
by: Barnett, Peter
Published: (2025)
by: Barnett, Peter
Published: (2025)
Will we run out of data? Limits of LLM scaling based on human-generated data
by: Villalobos, Pablo, et al.
Published: (2022)
by: Villalobos, Pablo, et al.
Published: (2022)
Frontier Large Language Models Rival State-of-the-Art Planners
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Exploring the Adversarial Capabilities of Large Language Models
by: Struppek, Lukas, et al.
Published: (2024)
by: Struppek, Lukas, et al.
Published: (2024)
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024)
by: Benton, Joe, et al.
Published: (2024)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Interactive Hyperparameter Optimization in Multi-Objective Problems via Preference Learning
by: Giovanelli, Joseph, et al.
Published: (2023)
by: Giovanelli, Joseph, et al.
Published: (2023)
Uncovering Capabilities of Model Pruning in Graph Contrastive Learning
by: Wu, Junran, et al.
Published: (2024)
by: Wu, Junran, et al.
Published: (2024)
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
by: Lin, Yong, et al.
Published: (2025)
by: Lin, Yong, et al.
Published: (2025)
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
by: Feng, Austin, et al.
Published: (2025)
by: Feng, Austin, et al.
Published: (2025)
Foundations and Frontiers of Graph Learning Theory
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
by: Miao, Tingjia, et al.
Published: (2026)
by: Miao, Tingjia, et al.
Published: (2026)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
In-context Learning of Evolving Data Streams with Tabular Foundational Models
by: Lourenço, Afonso, et al.
Published: (2025)
by: Lourenço, Afonso, et al.
Published: (2025)
Pre-trained Large Language Models Learn Hidden Markov Models In-context
by: Dai, Yijia, et al.
Published: (2025)
by: Dai, Yijia, et al.
Published: (2025)
You are out of context!
by: Cobino, Giancarlo, et al.
Published: (2024)
by: Cobino, Giancarlo, et al.
Published: (2024)
Overcoming Dependent Censoring in the Evaluation of Survival Models
by: Lillelund, Christian Marius, et al.
Published: (2025)
by: Lillelund, Christian Marius, et al.
Published: (2025)
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
by: Li, Tenghui, et al.
Published: (2024)
by: Li, Tenghui, et al.
Published: (2024)
Similar Items
-
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023) -
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024) -
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025) -
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025) -
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)