Shutdown Safety Valves for Advanced AI
Fuente:
arXiv
Saved in:
| Main Author: | Conitzer, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Holistic Safety and Responsibility Evaluations of Advanced AI Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Efficiently Solving Turn-Taking Stochastic Games with Extensive-Form Correlation
by: Zhang, Hanrui, et al.
Published: (2024)
by: Zhang, Hanrui, et al.
Published: (2024)
An Interpretable Automated Mechanism Design Framework with Large Language Models
by: Liu, Jiayuan, et al.
Published: (2025)
by: Liu, Jiayuan, et al.
Published: (2025)
Standardization Trends on Safety and Trustworthiness Technology for Advanced AI
by: Jeon, Jonghong
Published: (2024)
by: Jeon, Jonghong
Published: (2024)
NeuroAI for AI Safety
by: Mineault, Patrick, et al.
Published: (2024)
by: Mineault, Patrick, et al.
Published: (2024)
A Knowledge-Informed Large Language Model Framework for U.S. Nuclear Power Plant Shutdown Initiating Event Classification for Probabilistic Risk Assessment
by: Xian, Min, et al.
Published: (2024)
by: Xian, Min, et al.
Published: (2024)
Observation Interference in Partially Observable Assistance Games
by: Emmons, Scott, et al.
Published: (2024)
by: Emmons, Scott, et al.
Published: (2024)
Why should we ever automate moral decision making?
by: Conitzer, Vincent
Published: (2024)
by: Conitzer, Vincent
Published: (2024)
Towards Optimal Valve Prescription for Transcatheter Aortic Valve Replacement (TAVR) Surgery: A Machine Learning Approach
by: Paschalidis, Phevos, et al.
Published: (2025)
by: Paschalidis, Phevos, et al.
Published: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Transformer Vibration Forecasting for Advancing Rail Safety and Maintenance 4.0
by: Larese, Darío C., et al.
Published: (2025)
by: Larese, Darío C., et al.
Published: (2025)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
by: Liu, Chenruo, et al.
Published: (2025)
by: Liu, Chenruo, et al.
Published: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
by: Chen, Pin-Yu
Published: (2025)
by: Chen, Pin-Yu
Published: (2025)
Recent Advances in Generative AI for Healthcare Applications
by: Shokrollahi, Yasin, et al.
Published: (2023)
by: Shokrollahi, Yasin, et al.
Published: (2023)
Optimizing Fire Safety: Reducing False Alarms Using Advanced Machine Learning Techniques
by: Jamal, Muhammad Hassan, et al.
Published: (2025)
by: Jamal, Muhammad Hassan, et al.
Published: (2025)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Advancing Ocean State Estimation with efficient and scalable AI
by: Xiang, Yanfei, et al.
Published: (2025)
by: Xiang, Yanfei, et al.
Published: (2025)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024)
by: Putta, Pranav, et al.
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Generative AI Agents in Autonomous Machines: A Safety Perspective
by: Jabbour, Jason, et al.
Published: (2024)
by: Jabbour, Jason, et al.
Published: (2024)
From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems
by: Stefani, Thomas, et al.
Published: (2026)
by: Stefani, Thomas, et al.
Published: (2026)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
Harmonizing Human Insights and AI Precision: Hand in Hand for Advancing Knowledge Graph Task
by: Wang, Shurong, et al.
Published: (2024)
by: Wang, Shurong, et al.
Published: (2024)
Incentive-Aware Multi-Fidelity Optimization for Generative Advertising in Large Language Models
by: Liu, Jiayuan, et al.
Published: (2026)
by: Liu, Jiayuan, et al.
Published: (2026)
Safety Modulation: Enhancing Safety in Reinforcement Learning through Cost-Modulated Rewards
by: Zhang, Hanping, et al.
Published: (2025)
by: Zhang, Hanping, et al.
Published: (2025)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
by: Sholehrasa, Hossein, et al.
Published: (2025)
by: Sholehrasa, Hossein, et al.
Published: (2025)
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
by: Lu, Haoran, et al.
Published: (2025)
by: Lu, Haoran, et al.
Published: (2025)
Curriculum Learning for Safety Alignment
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
Reasoning as an Adaptive Defense for Safety
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
Speculative Safety-Aware Decoding
by: Wang, Xuekang, et al.
Published: (2025)
by: Wang, Xuekang, et al.
Published: (2025)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
by: Cho, Dongkyu Derek, et al.
Published: (2025)
by: Cho, Dongkyu Derek, et al.
Published: (2025)
Similar Items
-
Holistic Safety and Responsibility Evaluations of Advanced AI Models
by: Weidinger, Laura, et al.
Published: (2024) -
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025) -
Efficiently Solving Turn-Taking Stochastic Games with Extensive-Form Correlation
by: Zhang, Hanrui, et al.
Published: (2024) -
An Interpretable Automated Mechanism Design Framework with Large Language Models
by: Liu, Jiayuan, et al.
Published: (2025) -
Standardization Trends on Safety and Trustworthiness Technology for Advanced AI
by: Jeon, Jonghong
Published: (2024)