Understanding the First Wave of AI Safety Institutes: Characteristics, Functions, and Challenges
Fuente:
arXiv
Salvato in:
| Autori principali: | Araujo, Renan, Fort, Kristina, Guest, Oliver |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
di: Fort, Kristina
Pubblicazione: (2024)
di: Fort, Kristina
Pubblicazione: (2024)
Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis
di: Delaney, Oscar, et al.
Pubblicazione: (2024)
di: Delaney, Oscar, et al.
Pubblicazione: (2024)
Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence
di: Guest, Oliver, et al.
Pubblicazione: (2025)
di: Guest, Oliver, et al.
Pubblicazione: (2025)
Risk Reporting for Developers' Internal AI Model Use
di: Delaney, Oscar, et al.
Pubblicazione: (2026)
di: Delaney, Oscar, et al.
Pubblicazione: (2026)
Safety First: Psychological Safety as the Key to AI Transformation
di: Reich, Aaron, et al.
Pubblicazione: (2026)
di: Reich, Aaron, et al.
Pubblicazione: (2026)
Functional Misalignment in Human-AI Interactions on Digital Platforms
di: Lerman, Kristina
Pubblicazione: (2026)
di: Lerman, Kristina
Pubblicazione: (2026)
Towards decolonising computational sciences
di: Birhane, Abeba, et al.
Pubblicazione: (2020)
di: Birhane, Abeba, et al.
Pubblicazione: (2020)
Institutional AI: A Governance Framework for Distributional AGI Safety
di: Pierucci, Federico, et al.
Pubblicazione: (2026)
di: Pierucci, Federico, et al.
Pubblicazione: (2026)
Demonstrating Restraint
di: Patell, L. C. R., et al.
Pubblicazione: (2026)
di: Patell, L. C. R., et al.
Pubblicazione: (2026)
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
Benchmarking and Understanding Safety Risks in AI Character Platforms
di: Wei, Yiluo, et al.
Pubblicazione: (2025)
di: Wei, Yiluo, et al.
Pubblicazione: (2025)
AI Safety for Everyone
di: Gyevnar, Balint, et al.
Pubblicazione: (2025)
di: Gyevnar, Balint, et al.
Pubblicazione: (2025)
Report on the Conference on Ethical and Responsible Design in the National AI Institutes: A Summary of Challenges
di: Conklin, Sherri Lynn, et al.
Pubblicazione: (2024)
di: Conklin, Sherri Lynn, et al.
Pubblicazione: (2024)
How Should AI Safety Benchmarks Benchmark Safety?
di: Yu, Cheng, et al.
Pubblicazione: (2026)
di: Yu, Cheng, et al.
Pubblicazione: (2026)
AI Safety, Alignment, and Ethics (AI SAE)
di: Waldner, Dylan
Pubblicazione: (2025)
di: Waldner, Dylan
Pubblicazione: (2025)
Safety cases for frontier AI
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development
di: Šekrst, Kristina, et al.
Pubblicazione: (2024)
di: Šekrst, Kristina, et al.
Pubblicazione: (2024)
International AI Safety Report 2026
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
The BIG Argument for AI Safety Cases
di: Habli, Ibrahim, et al.
Pubblicazione: (2025)
di: Habli, Ibrahim, et al.
Pubblicazione: (2025)
Persuasion and Safety in the Era of Generative AI
di: Kong, Haein
Pubblicazione: (2025)
di: Kong, Haein
Pubblicazione: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
di: Dobbe, Roel
Pubblicazione: (2025)
di: Dobbe, Roel
Pubblicazione: (2025)
A Grading Rubric for AI Safety Frameworks
di: Alaga, Jide, et al.
Pubblicazione: (2024)
di: Alaga, Jide, et al.
Pubblicazione: (2024)
Evaluating AI Providers' Frontier Safety Frameworks
di: Stelling, Lily, et al.
Pubblicazione: (2025)
di: Stelling, Lily, et al.
Pubblicazione: (2025)
Astra: AI Safety, Trust, & Risk Assessment
di: Aggarwal, Pranav, et al.
Pubblicazione: (2026)
di: Aggarwal, Pranav, et al.
Pubblicazione: (2026)
Application of the Cyberinfrastructure Production Function Model to R1 Institutions
di: Smith, Preston M., et al.
Pubblicazione: (2025)
di: Smith, Preston M., et al.
Pubblicazione: (2025)
Navigating Fairness: Practitioners' Understanding, Challenges, and Strategies in AI/ML Development
di: Pant, Aastha, et al.
Pubblicazione: (2024)
di: Pant, Aastha, et al.
Pubblicazione: (2024)
AI Safety in Generative AI Large Language Models: A Survey
di: Chua, Jaymari, et al.
Pubblicazione: (2024)
di: Chua, Jaymari, et al.
Pubblicazione: (2024)
Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
di: Yew, Rui-Jie, et al.
Pubblicazione: (2025)
di: Yew, Rui-Jie, et al.
Pubblicazione: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
Aiming for AI Interoperability: Challenges and Opportunities
di: Faveri, Benjamin, et al.
Pubblicazione: (2026)
di: Faveri, Benjamin, et al.
Pubblicazione: (2026)
AI Alignment vs. AI Ethical Treatment: 10 Challenges
di: Bradley, Adam, et al.
Pubblicazione: (2025)
di: Bradley, Adam, et al.
Pubblicazione: (2025)
Information Retrieval Induced Safety Degradation in AI Agents
di: Yu, Cheng, et al.
Pubblicazione: (2025)
di: Yu, Cheng, et al.
Pubblicazione: (2025)
AI Safety Evaluations Need To Consider Cascading Effects
di: Neumann, Anna, et al.
Pubblicazione: (2026)
di: Neumann, Anna, et al.
Pubblicazione: (2026)
Assessing the Case for Africa-Centric AI Safety Evaluations
di: Ireri, Gathoni, et al.
Pubblicazione: (2026)
di: Ireri, Gathoni, et al.
Pubblicazione: (2026)
Beyond Model Readiness: Institutional Readiness for AI Deployment in Public Systems
di: Legara, Erika Fille, et al.
Pubblicazione: (2026)
di: Legara, Erika Fille, et al.
Pubblicazione: (2026)
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
Understanding Student Interaction with AI-Powered Next-Step Hints: Strategies and Challenges
di: Birillo, Anastasiia, et al.
Pubblicazione: (2025)
di: Birillo, Anastasiia, et al.
Pubblicazione: (2025)
Vernacularizing Taxonomies of Harm is Essential for Operationalizing Holistic AI Safety
di: Kennedy, Wm. Matthew, et al.
Pubblicazione: (2024)
di: Kennedy, Wm. Matthew, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
di: Fort, Kristina
Pubblicazione: (2024) -
Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis
di: Delaney, Oscar, et al.
Pubblicazione: (2024) -
Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence
di: Guest, Oliver, et al.
Pubblicazione: (2025) -
Risk Reporting for Developers' Internal AI Model Use
di: Delaney, Oscar, et al.
Pubblicazione: (2026) -
Safety First: Psychological Safety as the Key to AI Transformation
di: Reich, Aaron, et al.
Pubblicazione: (2026)