Saved in:
| Main Authors: | Stickland, Asa Cooper, Michelfeit, Jan, Mani, Arathi, Griffin, Charlie, Matthews, Ollie, Korbak, Tomek, Inglis, Rogan, Makins, Oliver, Cooney, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.13526 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025)
by: Lindner, David, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)
by: Stickland, Asa Cooper, et al.
Published: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Why Do Language Model Agents Whistleblow?
by: Agrawal, Kushal, et al.
Published: (2025)
by: Agrawal, Kushal, et al.
Published: (2025)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
Training Agents to Self-Report Misbehavior
by: Lee, Bruce W., et al.
Published: (2026)
by: Lee, Bruce W., et al.
Published: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
by: Han, Zhenyu, et al.
Published: (2025)
by: Han, Zhenyu, et al.
Published: (2025)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Reasoning Models Struggle to Control their Chains of Thought
by: Yueh-Han, Chen, et al.
Published: (2026)
by: Yueh-Han, Chen, et al.
Published: (2026)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
by: Bai, Huawei, et al.
Published: (2025)
by: Bai, Huawei, et al.
Published: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Propensity Inference: Environmental Contributors to LLM Behaviour
by: Järviniemi, Olli, et al.
Published: (2026)
by: Järviniemi, Olli, et al.
Published: (2026)
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
by: Hirose, Noriaki, et al.
Published: (2026)
by: Hirose, Noriaki, et al.
Published: (2026)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
AsyncDSB: Schedule-Asynchronous Diffusion Schrödinger Bridge for Image Inpainting
by: Han, Zihao, et al.
Published: (2024)
by: Han, Zihao, et al.
Published: (2024)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
by: Luo, Shuqing, et al.
Published: (2025)
by: Luo, Shuqing, et al.
Published: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
AsyncSC: An Asynchronous Sidechain for Multi-Domain Data Exchange in Internet of Things
by: Yang, Lingxiao, et al.
Published: (2024)
by: Yang, Lingxiao, et al.
Published: (2024)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
Surveillance as a Disciplinary Mechanism in Manjula Padmanabhan’s The Island of Lost Girls
by: Arathi Babu
Published: (2017)
by: Arathi Babu
Published: (2017)
AsyncBEV: Cross-modal Flow Alignment in Asynchronous 3D Object Detection
by: Wang, Shiming, et al.
Published: (2026)
by: Wang, Shiming, et al.
Published: (2026)
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios
by: Shi, Kou, et al.
Published: (2026)
by: Shi, Kou, et al.
Published: (2026)
AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
by: Ma, Lianjie, et al.
Published: (2026)
by: Ma, Lianjie, et al.
Published: (2026)
Safety case template for frontier AI: A cyber inability argument
by: Goemans, Arthur, et al.
Published: (2024)
by: Goemans, Arthur, et al.
Published: (2024)
AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation
by: Yang, Kai, et al.
Published: (2026)
by: Yang, Kai, et al.
Published: (2026)
AsyncEvGS: Asynchronous Event-Assisted Gaussian Splatting for Handheld Motion-Blurred Scenes
by: Dai, Jun, et al.
Published: (2026)
by: Dai, Jun, et al.
Published: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Aligning language models with human preferences
by: Korbak, Tomasz
Published: (2024)
by: Korbak, Tomasz
Published: (2024)
From Murder Mystery to Tragic Truth- The Cutthroat Revelation? A Case Report
by: S, Arathi, et al.
Published: (2025)
by: S, Arathi, et al.
Published: (2025)
Red novae, stellar mergers in binary and triple systems, and bipolar nebulae
by: Kaminski, Tomek
Published: (2024)
by: Kaminski, Tomek
Published: (2024)
Emergent Compositional Communication for Latent World Properties
by: Kaszyński, Tomek
Published: (2026)
by: Kaszyński, Tomek
Published: (2026)
Similar Items
-
Practical challenges of control monitoring in frontier AI deployments
by: Lindner, David, et al.
Published: (2025) -
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025) -
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024) -
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023) -
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)