Safety Cases: How to Justify the Safety of Advanced AI Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Clymer, Joshua, Gabrieli, Nick, Krueger, David, Larsen, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Affirmative safety: An approach to risk management for high-risk AI
von: Wasil, Akash R., et al.
Veröffentlicht: (2024)
von: Wasil, Akash R., et al.
Veröffentlicht: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
von: Dobbe, Roel
Veröffentlicht: (2025)
von: Dobbe, Roel
Veröffentlicht: (2025)
International Scientific Report on the Safety of Advanced AI (Interim Report)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
von: Petrova, Nora, et al.
Veröffentlicht: (2026)
von: Petrova, Nora, et al.
Veröffentlicht: (2026)
Open Problems in Machine Unlearning for AI Safety
von: Barez, Fazl, et al.
Veröffentlicht: (2025)
von: Barez, Fazl, et al.
Veröffentlicht: (2025)
The Elephant in the Room -- Why AI Safety Demands Diverse Teams
von: Rostcheck, David, et al.
Veröffentlicht: (2024)
von: Rostcheck, David, et al.
Veröffentlicht: (2024)
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
von: Barnett, Peter, et al.
Veröffentlicht: (2024)
von: Barnett, Peter, et al.
Veröffentlicht: (2024)
Toward an African Agenda for AI Safety
von: Segun, Samuel T., et al.
Veröffentlicht: (2025)
von: Segun, Samuel T., et al.
Veröffentlicht: (2025)
Concrete Problems in AI Safety, Revisited
von: Raji, Inioluwa Deborah, et al.
Veröffentlicht: (2023)
von: Raji, Inioluwa Deborah, et al.
Veröffentlicht: (2023)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
von: Chin, Yik Chan, et al.
Veröffentlicht: (2026)
von: Chin, Yik Chan, et al.
Veröffentlicht: (2026)
The Singapore Consensus on Global AI Safety Research Priorities
von: Bengio, Yoshua, et al.
Veröffentlicht: (2025)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2025)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
von: Staufer, Leon, et al.
Veröffentlicht: (2026)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
von: Xia, Boming, et al.
Veröffentlicht: (2024)
von: Xia, Boming, et al.
Veröffentlicht: (2024)
Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
von: Bajcsy, Andrea, et al.
Veröffentlicht: (2024)
von: Bajcsy, Andrea, et al.
Veröffentlicht: (2024)
AI Safety: Necessary, but insufficient and possibly problematic
von: P, Deepak
Veröffentlicht: (2024)
von: P, Deepak
Veröffentlicht: (2024)
Emerging Practices in Frontier AI Safety Frameworks
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
International AI Safety Report
von: Bengio, Yoshua, et al.
Veröffentlicht: (2025)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2025)
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
von: Shu, Wesley, et al.
Veröffentlicht: (2026)
von: Shu, Wesley, et al.
Veröffentlicht: (2026)
Justified Evidence Collection for Argument-based AI Fairness Assurance
von: Sabuncuoglu, Alpay, et al.
Veröffentlicht: (2025)
von: Sabuncuoglu, Alpay, et al.
Veröffentlicht: (2025)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
von: Feakins, Shaun, et al.
Veröffentlicht: (2026)
von: Feakins, Shaun, et al.
Veröffentlicht: (2026)
Upstream and Downstream AI Safety: Both on the Same River?
von: McDermid, John, et al.
Veröffentlicht: (2024)
von: McDermid, John, et al.
Veröffentlicht: (2024)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
von: Chiba-Okabe, Hiroaki
Veröffentlicht: (2024)
von: Chiba-Okabe, Hiroaki
Veröffentlicht: (2024)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
von: Costa, Mariana Lins
Veröffentlicht: (2026)
von: Costa, Mariana Lins
Veröffentlicht: (2026)
Combining Cost-Constrained Runtime Monitors for AI Safety
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
Building Effective Safety Guardrails in AI Education Tools
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
What Is AI Safety? What Do We Want It to Be?
von: Harding, Jacqueline, et al.
Veröffentlicht: (2025)
von: Harding, Jacqueline, et al.
Veröffentlicht: (2025)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
von: Han, Shanshan
Veröffentlicht: (2024)
von: Han, Shanshan
Veröffentlicht: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
von: Kempermann, Manon, et al.
Veröffentlicht: (2025)
von: Kempermann, Manon, et al.
Veröffentlicht: (2025)
Introduction to AI Safety, Ethics, and Society
von: Hendrycks, Dan
Veröffentlicht: (2024)
von: Hendrycks, Dan
Veröffentlicht: (2024)
Building Trust: Foundations of Security, Safety and Transparency in AI
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2024)
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2024)
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
von: Menke, Antonio-Gabriel Chacón, et al.
Veröffentlicht: (2025)
von: Menke, Antonio-Gabriel Chacón, et al.
Veröffentlicht: (2025)
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems
von: Olukola, Oluseyi, et al.
Veröffentlicht: (2026)
von: Olukola, Oluseyi, et al.
Veröffentlicht: (2026)
LLM Safety for Children
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
von: Reddy, Pavan, et al.
Veröffentlicht: (2025)
von: Reddy, Pavan, et al.
Veröffentlicht: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
von: Brophy, Matthew
Veröffentlicht: (2025)
von: Brophy, Matthew
Veröffentlicht: (2025)
Ähnliche Einträge
-
Affirmative safety: An approach to risk management for high-risk AI
von: Wasil, Akash R., et al.
Veröffentlicht: (2024) -
Safety Cases: A Scalable Approach to Frontier AI Safety
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025) -
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
von: Dobbe, Roel
Veröffentlicht: (2025) -
International Scientific Report on the Safety of Advanced AI (Interim Report)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024) -
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)