Saved in:
| Main Authors: | Griffin, Charlie, Thomson, Louis, Shlegeris, Buck, Abate, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.07985 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025)
by: Kutasov, Jon, et al.
Published: (2025)
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023)
by: Greenblatt, Ryan, et al.
Published: (2023)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
Efficiently Deploying LLMs with Controlled Risk
by: Zellinger, Michael J., et al.
Published: (2024)
by: Zellinger, Michael J., et al.
Published: (2024)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
Neural Proofs for Sound Verification and Control of Complex Systems
by: Abate, Alessandro
Published: (2025)
by: Abate, Alessandro
Published: (2025)
Networked Communication for Decentralised Agents in Mean-Field Games
by: Benjamin, Patrick, et al.
Published: (2023)
by: Benjamin, Patrick, et al.
Published: (2023)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
by: Jackermeier, Mathias, et al.
Published: (2024)
by: Jackermeier, Mathias, et al.
Published: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Networked Communication for Mean-Field Games with Function Approximation and Empirical Mean-Field Estimation
by: Benjamin, Patrick, et al.
Published: (2024)
by: Benjamin, Patrick, et al.
Published: (2024)
NeuroAI for AI Safety
by: Mineault, Patrick, et al.
Published: (2024)
by: Mineault, Patrick, et al.
Published: (2024)
Zero-Shot Instruction Following in RL via Structured LTL Representations
by: Giuri, Mattia, et al.
Published: (2025)
by: Giuri, Mattia, et al.
Published: (2025)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Resource-Efficient Generative AI Model Deployment in Mobile Edge Networks
by: Liang, Yuxin, et al.
Published: (2024)
by: Liang, Yuxin, et al.
Published: (2024)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
by: Su, Jianhao, et al.
Published: (2026)
by: Su, Jianhao, et al.
Published: (2026)
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024)
by: Benton, Joe, et al.
Published: (2024)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment
by: Digalakis Jr, Vassilis, et al.
Published: (2025)
by: Digalakis Jr, Vassilis, et al.
Published: (2025)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
by: Schnitzer, Yannik, et al.
Published: (2026)
by: Schnitzer, Yannik, et al.
Published: (2026)
Zero-Shot Instruction Following in RL via Structured LTL Representations
by: Jackermeier, Mathias, et al.
Published: (2026)
by: Jackermeier, Mathias, et al.
Published: (2026)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Evaluating Machine Learning Models against Clinical Protocols for Enhanced Interpretability and Continuity of Care
by: Sirocchi, Christel, et al.
Published: (2024)
by: Sirocchi, Christel, et al.
Published: (2024)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
by: Prinster, Drew, et al.
Published: (2025)
by: Prinster, Drew, et al.
Published: (2025)
Shutdown Safety Valves for Advanced AI
by: Conitzer, Vincent
Published: (2026)
by: Conitzer, Vincent
Published: (2026)
AI-CARE: Carbon-Aware Reporting Evaluation Metric for AI Models
by: Santosh, KC, et al.
Published: (2026)
by: Santosh, KC, et al.
Published: (2026)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
by: Almurshed, Osama, et al.
Published: (2025)
by: Almurshed, Osama, et al.
Published: (2025)
Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment
by: Wang, Xubin, et al.
Published: (2025)
by: Wang, Xubin, et al.
Published: (2025)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Studying Cross-cluster Modularity in Neural Networks
by: Golechha, Satvik, et al.
Published: (2025)
by: Golechha, Satvik, et al.
Published: (2025)
Similar Items
-
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024) -
Evaluating Control Protocols for Untrusted AI Agents
by: Kutasov, Jon, et al.
Published: (2025) -
AI Control: Improving Safety Despite Intentional Subversion
by: Greenblatt, Ryan, et al.
Published: (2023) -
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022) -
Efficiently Deploying LLMs with Controlled Risk
by: Zellinger, Michael J., et al.
Published: (2024)