Ctrl-Z: Controlling AI Agents via Resampling
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatt, Aryan, Rushing, Cody, Kaufman, Adam, Tracy, Tyler, Georgiev, Vasil, Matolcsi, David, Khan, Akbir, Shlegeris, Buck |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025)
by: Kaufman, Adam, et al.
Published: (2025)
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023)
by: Greenblatt, Ryan, et al.
Published: (2023)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Explorations of Self-Repair in Language Models
by: Rushing, Cody, et al.
Published: (2024)
by: Rushing, Cody, et al.
Published: (2024)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025)
by: Kutasov, Jon, et al.
Published: (2025)
Basic Legibility Protocols Improve Trusted Monitoring
by: Sreevatsa, Ashwin, et al.
Published: (2026)
by: Sreevatsa, Ashwin, et al.
Published: (2026)
Retrying vs Resampling in AI Control
by: Lucassen, James, et al.
Published: (2026)
by: Lucassen, James, et al.
Published: (2026)
Factorio Learning Environment
by: Hopkins, Jack, et al.
Published: (2025)
by: Hopkins, Jack, et al.
Published: (2025)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Controllable Generation via Locally Constrained Resampling
by: Ahmed, Kareem, et al.
Published: (2024)
by: Ahmed, Kareem, et al.
Published: (2024)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026)
by: Gan, Eric, et al.
Published: (2026)
Peirce in the Machine: How Mixture of Experts Models Perform Hypothesis Construction
by: Rushing, Bruce
Published: (2024)
by: Rushing, Bruce
Published: (2024)
Interpolating Discrete Diffusion Models with Controllable Resampling
by: Kollovieh, Marcel, et al.
Published: (2026)
by: Kollovieh, Marcel, et al.
Published: (2026)
Variance Reduction via Resampling and Experience Replay
by: Han, Jiale, et al.
Published: (2025)
by: Han, Jiale, et al.
Published: (2025)
Ctrl-A: Control-Driven Online Data Augmentation
by: Christensen, Jesper B., et al.
Published: (2026)
by: Christensen, Jesper B., et al.
Published: (2026)
Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL
by: Chen, Xingyu, et al.
Published: (2025)
by: Chen, Xingyu, et al.
Published: (2025)
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026)
by: Cheng, Emily, et al.
Published: (2026)
Explainable AI Approach using Near Misses Analysis
by: Kaufman, Eran, et al.
Published: (2024)
by: Kaufman, Eran, et al.
Published: (2024)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
An Operator-Consistent Graph Neural Network for Learning Diffusion Dynamics on Irregular Meshes
by: Li, Yuelian, et al.
Published: (2025)
by: Li, Yuelian, et al.
Published: (2025)
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024)
by: Lin, Kuan Heng, et al.
Published: (2024)
Interleaved Resampling and Refitting: Data and Compute-Efficient Evaluation of Black-Box Predictors
by: Hu, Haichen, et al.
Published: (2026)
by: Hu, Haichen, et al.
Published: (2026)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
by: Jenner, Erik, et al.
Published: (2024)
by: Jenner, Erik, et al.
Published: (2024)
Resampling and averaging coordinates on data
by: Blumberg, Andrew J., et al.
Published: (2024)
by: Blumberg, Andrew J., et al.
Published: (2024)
Ctrl-GenAug: Controllable Generative Augmentation for Medical Sequence Classification
by: Zhou, Xinrui, et al.
Published: (2024)
by: Zhou, Xinrui, et al.
Published: (2024)
Reachability Weighted Offline Goal-conditioned Resampling
by: Yang, Wenyan, et al.
Published: (2025)
by: Yang, Wenyan, et al.
Published: (2025)
Analyzing Deep Transformer Models for Time Series Forecasting via Manifold Learning
by: Kaufman, Ilya, et al.
Published: (2024)
by: Kaufman, Ilya, et al.
Published: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
by: Sandoval, Aaron, et al.
Published: (2025)
by: Sandoval, Aaron, et al.
Published: (2025)
The Limits of Inference Scaling Through Resampling
by: Stroebl, Benedikt, et al.
Published: (2024)
by: Stroebl, Benedikt, et al.
Published: (2024)
Decoupled Prioritized Resampling for Offline RL
by: Yue, Yang, et al.
Published: (2023)
by: Yue, Yang, et al.
Published: (2023)
Resampling-free Particle Filters in High-dimensions
by: Boopathy, Akhilan, et al.
Published: (2024)
by: Boopathy, Akhilan, et al.
Published: (2024)
ButterflyMoE: Sub-Linear Ternary Experts via Structured Butterfly Orbits
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
by: Cook, Jonathan, et al.
Published: (2025)
by: Cook, Jonathan, et al.
Published: (2025)
Neural Bipartite Matching
by: Georgiev, Dobrik, et al.
Published: (2020)
by: Georgiev, Dobrik, et al.
Published: (2020)
Similar Items
-
BashArena: A Control Setting for Highly Privileged AI Agents
by: Kaufman, Adam, et al.
Published: (2025) -
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023) -
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024) -
Explorations of Self-Repair in Language Models
by: Rushing, Cody, et al.
Published: (2024) -
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)