The BIG Argument for AI Safety Cases
Fuente:
arXiv
Salvato in:
| Autori principali: | Habli, Ibrahim, Hawkins, Richard, Paterson, Colin, Ryan, Philippa, Jia, Yan, Sujan, Mark, McDermid, John |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Upstream and Downstream AI Safety: Both on the Same River?
di: McDermid, John, et al.
Pubblicazione: (2024)
di: McDermid, John, et al.
Pubblicazione: (2024)
What's my role? Modelling responsibility for AI-based safety-critical systems
di: Ryan, Philippa, et al.
Pubblicazione: (2023)
di: Ryan, Philippa, et al.
Pubblicazione: (2023)
Unravelling Responsibility for AI
di: Porter, Zoe, et al.
Pubblicazione: (2023)
di: Porter, Zoe, et al.
Pubblicazione: (2023)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
di: Feakins, Shaun, et al.
Pubblicazione: (2026)
di: Feakins, Shaun, et al.
Pubblicazione: (2026)
The case for delegated AI autonomy for Human AI teaming in healthcare
di: Jia, Yan, et al.
Pubblicazione: (2025)
di: Jia, Yan, et al.
Pubblicazione: (2025)
Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems
di: Hodge, Victoria J., et al.
Pubblicazione: (2025)
di: Hodge, Victoria J., et al.
Pubblicazione: (2025)
Fair by design: A sociotechnical approach to justifying the fairness of AI-enabled systems across the lifecycle
di: Kaas, Marten H. L., et al.
Pubblicazione: (2024)
di: Kaas, Marten H. L., et al.
Pubblicazione: (2024)
Learning Run-time Safety Monitors for Machine Learning Components
di: Vardal, Ozan, et al.
Pubblicazione: (2024)
di: Vardal, Ozan, et al.
Pubblicazione: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
Insights from Railway Professionals: Rethinking Railway assumptions regarding safety and autonomy
di: Hunter, Josh, et al.
Pubblicazione: (2025)
di: Hunter, Josh, et al.
Pubblicazione: (2025)
Safety Analysis of Autonomous Railway Systems: An Introduction to the SACRED Methodology
di: Hunter, Josh, et al.
Pubblicazione: (2024)
di: Hunter, Josh, et al.
Pubblicazione: (2024)
Assessing the Case for Africa-Centric AI Safety Evaluations
di: Ireri, Gathoni, et al.
Pubblicazione: (2026)
di: Ireri, Gathoni, et al.
Pubblicazione: (2026)
International AI Safety Report 2026
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
International Scientific Report on the Safety of Advanced AI (Interim Report)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
Milestone Determination for Autonomous Railway Operation
di: Hunter, Josh, et al.
Pubblicazione: (2025)
di: Hunter, Josh, et al.
Pubblicazione: (2025)
Arguing conformance with data protection principles
di: Smith, Chris, et al.
Pubblicazione: (2026)
di: Smith, Chris, et al.
Pubblicazione: (2026)
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
Examining Popular Arguments Against AI Existential Risk: A Philosophical Analysis
di: Swoboda, Torben, et al.
Pubblicazione: (2025)
di: Swoboda, Torben, et al.
Pubblicazione: (2025)
Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
di: Yan, Lixiang, et al.
Pubblicazione: (2025)
di: Yan, Lixiang, et al.
Pubblicazione: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
AI Safety for Everyone
di: Gyevnar, Balint, et al.
Pubblicazione: (2025)
di: Gyevnar, Balint, et al.
Pubblicazione: (2025)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
di: Fort, Kristina
Pubblicazione: (2024)
di: Fort, Kristina
Pubblicazione: (2024)
Safety First: Psychological Safety as the Key to AI Transformation
di: Reich, Aaron, et al.
Pubblicazione: (2026)
di: Reich, Aaron, et al.
Pubblicazione: (2026)
How Should AI Safety Benchmarks Benchmark Safety?
di: Yu, Cheng, et al.
Pubblicazione: (2026)
di: Yu, Cheng, et al.
Pubblicazione: (2026)
Building Trust: Foundations of Security, Safety and Transparency in AI
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
AI Safety, Alignment, and Ethics (AI SAE)
di: Waldner, Dylan
Pubblicazione: (2025)
di: Waldner, Dylan
Pubblicazione: (2025)
International AI Safety Report
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
Safety cases for frontier AI
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
Justified Evidence Collection for Argument-based AI Fairness Assurance
di: Sabuncuoglu, Alpay, et al.
Pubblicazione: (2025)
di: Sabuncuoglu, Alpay, et al.
Pubblicazione: (2025)
Persuasion and Safety in the Era of Generative AI
di: Kong, Haein
Pubblicazione: (2025)
di: Kong, Haein
Pubblicazione: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
di: Brundage, Miles, et al.
Pubblicazione: (2026)
di: Brundage, Miles, et al.
Pubblicazione: (2026)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
Divergent Paths to Depolarization: Dialogue Design Determines the Prosocial Benefits of AI-Assisted Political Argumentation
di: Zhu, Jianlong, et al.
Pubblicazione: (2026)
di: Zhu, Jianlong, et al.
Pubblicazione: (2026)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
di: Dobbe, Roel
Pubblicazione: (2025)
di: Dobbe, Roel
Pubblicazione: (2025)
Evaluating AI Providers' Frontier Safety Frameworks
di: Stelling, Lily, et al.
Pubblicazione: (2025)
di: Stelling, Lily, et al.
Pubblicazione: (2025)
A Grading Rubric for AI Safety Frameworks
di: Alaga, Jide, et al.
Pubblicazione: (2024)
di: Alaga, Jide, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Upstream and Downstream AI Safety: Both on the Same River?
di: McDermid, John, et al.
Pubblicazione: (2024) -
What's my role? Modelling responsibility for AI-based safety-critical systems
di: Ryan, Philippa, et al.
Pubblicazione: (2023) -
Unravelling Responsibility for AI
di: Porter, Zoe, et al.
Pubblicazione: (2023) -
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
di: Feakins, Shaun, et al.
Pubblicazione: (2026) -
The case for delegated AI autonomy for Human AI teaming in healthcare
di: Jia, Yan, et al.
Pubblicazione: (2025)